SOURCE-LINKED INTELLIGENCE
Learning Representations through Token Prediction: Geometry, Approximation, and Downstream Guarantees
Token prediction is a central pre-training objective for modern language models. Despite its empirical success, why token prediction learns broadly useful representations remains incompletely understood. We develop a statistical framework connecting token prediction with representation geometry, encoder approximation, and downstream performance. Under a softmax prediction head, we show that accurate token prediction organizes token embeddings according to similarities between the distributions of contexts in which different token types appear, as measured by Hellinger distance, with explicit e
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-08-30T22:33:15.000Z
First collected: 2026-09-21T07:22:03.933Z. This is not the publication date.