SOURCE-LINKED INTELLIGENCE
One Spectrum, Two Resources: Data-Memory Scaling in Autoregressive Prediction
How much learned memory is needed to benefit from more data? We show that the two resources are governed by one predictive-energy spectrum in a positive-entropy autoregressive retrieval source. Each coordinate contributes its query probability times the squared radius of its unknown logit. Writing $μ$ for the resulting energy spectrum, we prove the minimax law $\mathfrak R^*_{\rm value}(n,B)\asymp_R Φ_μ(n^{-1})+Φ_μ(τ_B), Φ_μ(t)=\int\min\{x,t\}\,μ(\mathrm dx),$ for $n$ prediction blocks and a learned state with at most $2^B$ values. Data set the resolution $1/n$; memory sets the level $τ_B$ rea
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-11T20:05:13.000Z
First collected: 2026-09-20T16:41:15.630Z. This is not the publication date.