Legibility is Not Interpretability: Evaluating Judged Importance versus Actual Importance in Chain-Of-Thought Reasoning Steps
Abstract
Reasoning traces from chain-of-thought models appear to offer a legible window into how a model arrives at its answer. A growing body of work treats them as such, using LLM judges to diagnose errors, evaluate faithfulness, and provide step-level supervision via process reward models and generative critics. These practices rely on the text of a reasoning step carrying information about its functional role. But does the text actually encode information about which reasoning steps matter? We operationalize the importance of a reasoning step as its advantage: the change in expected reward, e.g., producing the correct final answer, from including that step, estimated via Monte Carlo rollouts. Using these estimates as ground truth, we evaluate whether LLM judges can identify high-advantage steps and find that they perform poorly across multiple models and datasets, including in controlled settings where cue-based manipulations create clear expectations about which steps should matter. Fine-tuning a model as a step-level critic yields meaningful improvement but remains distant from the gold labels, suggesting that step importance is only partially recoverable from the text of the reasoning trace. Our findings contribute to a growing body of work on chain-of-thought faithfulness that cautions against treating the legibility of reasoning traces as interpretability and raise questions about the validity of judge-based analyses of model reasoning.