TRACES: Tagging Reasoning Steps for Adaptive Cost-Efficient Early-Stopping
Yannis Belkhiter ⋅ Seshu Tirupathi ⋅ Giulio Zizzo ⋅ John Kelleher
Abstract
The field of Language Reasoning Models (LRMs) has been very active over the past few years with advances in training and inference techniques enabling LRMs to reason longer, and more accurately. However, a growing body of studies show that LRMs are still inefficient, over-generating verification and reflection steps. Additionally, the high-level role of each reasoning steps and how these different step types contributes to the generation of correct answers, is largely underexplored. To address this challenge, we introduce TRACES: a lightweight framework for online $\textbf{T}$agging of the $\textbf{R}$easoning steps enabling $\textbf{A}$daptive $\textbf{C}$ost-$\textbf{E}$fficient Early-$\textbf{S}$topping of LRM inferences. Building on this framework we monitor reasoning behaviors during inferences, and we find that LRMs tend to shift their reasoning behavior after reaching a correct answer. We demonstrate that the monitoring of the specific type of steps can produce effective interpretable early stopping criteria. We evaluate the TRACES framework on three mathematical reasoning benchmarks, namely, MATH500, GSM8K, AIME and two knowledge and reasoning benchmarks, MMLU and GPQA respectively. We achieve 20 to 50\% token reduction while maintaining comparable accuracy to standard generation.
Successful Page Load