SOURCE-LINKED INTELLIGENCE
Refuse, Decompose, Refresh: A Claim-Safe Protocol for Closed-Loop AI Evaluation
An AI evaluation can be perfectly reproducible and still support the wrong claim. This risk is acute in closed-loop systems: policy determines visited states, observable components, and which failures leave a measurable trace. We propose a claim-safe protocol with three actions. Refuse: abstain when a clean reference stream or matched runtime comparison lacks support. Decompose: report protocol execution, operational false admission, and structural hypotheses separately rather than as one PASS/FAIL label. Refresh: treat distribution-shift alarms as requests to invalidate and recompute a refere
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-17T15:09:16.000Z
- arXiv · Artificial Intelligence · 2026-09-17T15:09:16.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.