Understanding and Mitigating Premature Confidence for Better LLM Reasoning
Jingchu Gai ⋅ Guanning Zeng ⋅ Christina Baek ⋅ Chen Wu ⋅ J Kolter ⋅ Andrej Risteski ⋅ Aditi Raghunathan
Abstract
It is commonly believed that increasing test-time compute improves performance, and that longer chain-of-thought (CoT) reasoning enables models to answer questions more effectively. However, we observe that models often exhibit \emph{premature confidence} -- they commit to an answer with high confidence early in its reasoning, rather than progressively build up confidence with more generated tokens. We show that premature confidence strongly correlates with logical shortcuts in CoT, establishing it as a scalable, annotation-free proxy for detecting low-quality reasoning. Building on this finding, we propose a progressive confidence shaping that penalizes premature-confident reasoning during reinforcement learning, encouraging models to produce CoT where confidence builds progressively. Our method consistently improves both accuracy and reasoning quality across synthetic reasoning (Countdown), mathematical reasoning (DAPO, AIME), and scientific reasoning (ScienceQA) benchmarks at scales from 1.5B to 8B parameters---on Countdown, accuracy improves by $3.2\times$ (+42.0pp) while the logical shortcut proportion drops by 48pp; on AIME, Pass@64 improves by 6.6pp. Our method also outperforms vanilla RL on a safety benchmark, producing models that more transparently surface misleading content in their reasoning traces. Finally, we conduct controlled experiments to investigate when and why premature confidence arises, and show when and why mitigating it yields larger improvements on harder problems and bigger models.
Successful Page Load