Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning
Heyang Jiang ⋅ Henry Liu ⋅ Baharan Mirzasoleiman
Abstract
Reinforcement learning with verifiable rewards (RLVR) has emerged as a highly effective framework for improving LLM reasoning, with methods such as GRPO among its most successful instantiations. However, GRPO relies on repeated generation of long chain-of-thought rollouts. Training time scales with the number of rollouts, a large fraction of which are uninformative. Thus, GRPO is computationally expensive and unstable. To mitigate this, existing approaches either generate a larger pool of rollouts and filter the most informative prompts, or leverage historical signals for filtering at later stages of training. These strategies offer modest performance gains, but slow down the overall process. To address this, we propose VarIance Guided Online Rollout allocation (VIGOR) which instead of allocating a fixed rollout budget per example, begins with a small number of rollouts for all examples in a batch and iteratively allocates additional rollouts to those with the highest group reward variance until a fixed total rollout budget is reached. Theoretically, we show that under the binary-reward setting of GRPO, within-group reward variance directly controls the gradient magnitude and speeds up training. Experiments on three model scales show that VIGOR reaches the target accuracy with up to 3.6$\times$ fewer rollouts, outperforms GRPO and baselines by up to 2.4\%, and further improves training stability and final performance under extended training.
Successful Page Load