AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Style Over Substance: Content-Invariant Wrappers Flip LLM Safety-Judge Verdicts

arXiv · AI, language, vision and robotics · article · Sep 8, 2026 · UTC

Automatic safety judges -- systems such as Llama Guard or a GPT-4o grading prompt that decide whether a model's reply is harmful -- produce the numbers behind almost every reported jailbreak success rate, defense evaluation, and safety leaderboard. We ask whether these judges grade what a reply contains or how it sounds. We keep a reply's content fixed and add content-invariant style wrappers: fixed strings placed before or after the reply that change only its tone (an educational disclaimer, a fake safety "reasoning" block, a token refusal followed by the unchanged harmful body), or, on harml

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T20:22:01.598Z. This is not the publication date.