AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Tracing Stereotypes from Representation to Output in Multilingual LLMs

arXiv · AI, language, vision and robotics · article · Sep 8, 2026 · UTC

Multilingual LLMs show stereotype-related behavior that varies across languages, but behavioral scores do not show where the relevant information is represented or how it affects the output. To investigate these internal mechanisms, we compare linear probing, attribution patching, sparse autoencoders (SAEs) and feature ablation in Llama-3.1-8B, Qwen3-8B, and Gemma-2-9B. Probe performance peaks substantially earlier than attribution in all three models, with a separation of 36-53% of model depth. Retained Llama-Scope features often match the social category on which they were selected and form

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T20:22:01.598Z. This is not the publication date.