SOURCE-LINKED INTELLIGENCE
Double descent is the principle of least action
The test error of a model plotted against its number of parameters $d$ falls, peaks when the model can just fit the training data, and falls again, exhibiting the double descent phenomenon. We explain the phenomenon with statistical mechanics. The training trajectory of a stochastic gradient-based method is a particle wandering over the energy landscape of the training loss at an induced temperature $T$, and a run that has equilibrated visits every parameter vector of a given training loss equally often, the fundamental postulate of statistical mechanics, with probability given by the Boltzman
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-16T17:17:46.000Z
- arXiv · Artificial Intelligence · 2026-09-16T17:17:46.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.