Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models
Jean de Dieu Nyandwi ⋅ Leena Mathur ⋅ Yonatan Bisk ⋅ Graham Neubig
Abstract
In reasoning models, concretely which reasoning behaviors improve the final accuracy on reasoning tasks? And if particular behaviors are more effective, are these the behaviors that are actually amplified by reasoning-based training? The answer to this question is consequential, because if the answer to the latter question is no, then it potentially reveals a significant inefficiency on how we are training models to reason. To help answer these questions, in this paper we develop an analysis framework based on behavioral lift, a concrete metric of how much a particular behavior improves (or degrades) reasoning model performance. Based on a behavioral lift analysis of 15 models across 6 benchmarks spanning both text-only and vision-language models, we demonstrate that several behaviors are indicative of success on reasoning tasks -- confidence calibration, knowledge alignment, and self-awareness. Simultaneously, several behaviors -- uncertainty acknowledgment, hypothesis testing, and self-correction -- yield minimal or even negative results. However, surprisingly, we find that it is not the case that RL-based training of reasoning models amplifies the most effective behaviors. For instance, uncertainty acknowledgment is amplified by 5--7$\times$ despite being only barely indicative of success, while confidence calibration is barely amplified, despite its utility. This points to future opportunities to guide reasoning-model training toward more effective behaviors, rather than relying solely on task success as a signal.
Successful Page Load