SOURCE-LINKED INTELLIGENCE
Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents
To keep coding agents from going off the rails, production systems now review each proposed action with a blocking monitor that can reject it before it runs (Auto Mode in Claude Code, Guardian in OpenAI's Codex). Prior evaluations of such monitors largely measure robustness to accidental harm or prompt injections from untrusted sources looking to hijack the agent. Less understood is how they hold up when the agent they monitor is persistently misaligned. To understand this risk, we task an adversarial agent with evading production blocking monitors and causing catastrophic harm, e.g. by exfilt
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-17T02:14:53.000Z
- arXiv · Artificial Intelligence · 2026-09-17T02:14:53.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.