AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens

arXiv · AI, language, vision and robotics · article · Sep 1, 2026 · UTC

A language model's prediction of its next token develops across layers, and lens methods track this process by decoding intermediate hidden states into tokens. But a lens reading reflects both the hidden state and the readout (the unembedding matrix) used to decode it. Many lenses are fit on a corpus, and we show that two lenses differing only in their fitting corpus can report different tokens for the same hidden states. We call this dependence corpus conditionality. To examine readout structure independently of the fitting corpus, we introduce Sparse Readout Prism (SRP), which decomposes the

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T05:51:54.566Z. This is not the publication date.