SOURCE-LINKED INTELLIGENCE
How do World Models and Policies Compose in LLM Agents? A Joint Spectral and Behavioral Account
How do LLM agents come to both understand environments they act in and master tasks set within them? Through controlled experiments combining world-model training (next-state prediction) and policy training (reward maximization), we investigate this question. We dissect the resulting models through their additive parameter updates. Geometrically, we find effective world-model updates are low-rank and share an input-feature subspace with policy updates while writing to nearly orthogonal output directions, whether trained separately or sequentially. However, we find that, in projection intervent
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-08-30T22:21:16.000Z
First collected: 2026-09-21T07:22:03.933Z. This is not the publication date.