Learn from Zero: Policy Optimization under Vanishing Advantage
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a vital paradigm for augmenting the reasoning capabilities of Large Language Models (LLMs). However, existing methods such as Group Relative Policy Optimization (GRPO) encounter critical limitations when sampled responses yield identical rewards. This phenomenon leads to vanishing gradient, effectively discarding valuable training signals from zero-variance prompts and thereby constraining both exploratory diversity and reasoning proficiency. To address this issue, we propose LZPO, a novel algorithm specifically engineered to recover training signals from zero-variance scenarios. LZPO introduces a semantic-entropy–guided advantage mechanism, which adaptively incentivizes correct reasoning paths while penalizing incorrect ones, even in the absence of reward variance. Furthermore, for prompts resulting in uniform failure, LZPO incorporates a reflective mechanism that guides the model to analyze errors and regenerate responses. By assigning advantages to latent reasoning trajectories based on semantic coherence, our approach effectively revitalizes exploratory learning signals. Experiments on six mathematical reasoning benchmarks show that LZPO achieves gains of up to 3.3 points over GRPO and outperforms other baselines, with particularly pronounced improvements on challenging reasoning tasks. Code is available at https://anonymous.4open.science/r/LZPO_RL.