Start Classifying: Categorical Critics for LLM Reinforcement Learning
Zhijian Zhou ⋅ Long Li ⋅ Xuan Zhang ⋅ Zongkai Liu ⋅ Yulei Qin ⋅ Ke Li ⋅ Xing Sun ⋅ Xiaoyu Tan ⋅ Chao Qu ⋅ Yuan Qi
Abstract
Proximal Policy Optimization (PPO) for large language models typically trains the critic with mean-squared-error (MSE) regression against scalar value targets. In reinforcement learning with verifiable rewards (RLVR), rewards are sparse and binary ($r \in \{0, 1\}$). Consequently, for intermediate reasoning states, the true return distribution is inherently bimodal, representing the bifurcated probabilities of eventual success or failure. We argue that standard MSE critics suffer from a distributional bottleneck: by collapsing these discrete, bimodal outcomes into a single scalar mean, MSE provides an impoverished supervision signal that distorts advantage estimation and restricts the critic's representational capacity. We propose HL-Gauss PPO, a drop-in replacement that trains the critic as a categorical predictor over a discretized value support using cross-entropy, while leaving the actor objective unchanged. Unlike prior work that employs categorical critics to stabilize bootstrapping in off-policy RL, we demonstrate that in on-policy RLVR, classification-based critics are superior because they natively match the bimodal nature of the reward landscape. Across mathematical reasoning and agentic Search-R1 benchmarks, HL-Gauss PPO consistently improves over strong PPO-based baselines and DAPO. Our analysis shows that categorical critics yield more symmetric, lower-variance advantage signals, providing a more robust training signal for policy updates.
Successful Page Load