SOURCE-LINKED INTELLIGENCE
Unifying ICL, SFT, KL-Regularized RL Through a Bayesian Lens
Supervised fine-tuning (SFT), few-shot in-context learning (ICL), KL-regularized RLHF/RLVR, and on-policy distillation are usually treated as distinct post-training paradigms. We develop a unified Bayesian perspective in which each is an instance of a two-step template: construct a (generalized) Bayes or Gibbs posterior from a reference model and a utility signal (log-likelihood, reward, or advantage), then approximate it by a forward-KL projection onto a parametric family, either in-weights (SFT/RL) or in-context (ICL). This yields a single chain of equivalences: few-shot ICL is an amortized
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-04T13:06:29.000Z
First collected: 2026-09-20T21:52:07.471Z. This is not the publication date.