AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Using OCR Heads to Verbalize Image Semantics

arXiv · AI, language, vision and robotics · article · Sep 16, 2026 · UTC

How do VLMs map from pixels to semantics? To understand this general question, we focus on a narrow one: studying how VLMs perform optical character recognition (OCR). Across four models, we identify attention heads causally necessary for OCR, and discover that these are in fact general-purpose heads that output interpretable semantic features across all image tokens. For example, pointing these heads at an image token containing the word "bike" causes Qwen3-VL-8B to output "bike," but pointing them at a bird wing causes the model to output the token "feathers." We collapse these heads' attent

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-19T20:28:26.698Z. This is not the publication date.