SOURCE-LINKED INTELLIGENCE
Quantifying the Memorization-to-Generalization Transition: Scaling Laws and Phase Structure in Grokking
Neural networks trained past memorization frequently undergo a delayed transition to generalization, a phenomenon known as grokking. Despite theoretical progress on \emph{why} this transition occurs, the quantitative structure of \emph{when} it occurs in hyperparameter space remains uncharacterized. We map the memorization-to-generalization boundary across 384 configurations of two-hidden-layer MLPs on modular arithmetic, fitting a power-law scaling relation for generalization onset time: $T_{\mathrm{grok}} \propto H^{-0.27}\, D^{-2.04}\, η^{-0.50}\, λ^{-0.64}$ ($R^2 = 0.732$; $0.821$ with int
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-09T16:47:03.000Z
First collected: 2026-09-20T19:32:24.350Z. This is not the publication date.