AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

You Can't Escape Your Own Activations : Evaluation Awareness and Multi-Agent Monitoring

arXiv · AI, language, vision and robotics · article · Sep 2, 2026 · UTC

LLM agents are increasingly deployed in multi-agent systems, where they can collude while keeping their actions benign. Output monitors designed to detect such collusions can be fooled by obfuscation and steganography, motivating the use of probes trained on internal activations. However, these probes are usually evaluated on agents that do not know they are being watched. We study how activation-based detection changes when agents are explicitly informed that their internal activations are being monitored, and when they additionally receive feedback from the monitor. We keep the models, probe

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T05:11:56.580Z. This is not the publication date.