AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Semantic Fibers and Cross-Gram Interference: A Calculus of Safety Drift in Overcomplete Representations

arXiv · AI, language, vision and robotics · article · Sep 14, 2026 · UTC

A deployed language model may refuse a harmful request in English yet comply with its faithful translation, revealing a cross-lingual safety failure that cannot be characterized reliably by output behavior alone. We formalize this phenomenon through an audited equivalence relation and show that, for a declared quotient, representation, metric, feature dictionary, scoring head, threshold, and contrast model, the resulting safety drift admits an exact linear-algebraic characterization. Specifically, the drift is a cross-Gram functional of the within-fiber contrast; its worst admissible value is

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T12:21:05.240Z. This is not the publication date.