AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs

arXiv · AI, language, vision and robotics · article · Sep 1, 2026 · UTC

How to divide a fixed annotation budget between supervised fine-tuning (SFT) and reinforcement learning (RL) during LLM post-training remains an open problem. Existing work characterizes only broad trends (e.g., SFT dominates in low-data regimes), lacks a principled allocation framework, and does not examine whether the optimal ratio transfers across model sizes. We frame this problem in terms of near-optimality: rather than seeking a single optimal SFT-RL ratio, we characterize the near-optimal region, the set of allocations within a specified tolerance of peak performance. Empirically, this

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T06:01:56.170Z. This is not the publication date.