SOURCE-LINKED INTELLIGENCE
AUDITPLAN: Commit, Then Answer for Auditable Safety Alignment
Safety tuning pipelines judge only the final answer, which makes it difficult to distinguish robust refusal from two undesirable shortcuts: blanket refusal on benign requests and polished but unfaithful safety rationales that do not actually constrain the answer. We propose AUDITPLAN, a single-model plan-then-answer approach where the model first emits a compact structured safety plan and then answers conditioned on it. The plan records a threat label, intended action, and explicit constraints, enabling machine-checkable auditing while remaining hidden from users at deployment. We train this b
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-16T18:45:24.000Z
- arXiv · Artificial Intelligence · 2026-09-16T18:45:24.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.