Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling
Abstract
Token is the fundamental unit of computation in modern autoregressive models, directly determining both inference cost and reasoning performance. Despite its importance, existing approaches to length control and prediction operate primarily at the coarse-grained sequence level. In this paper, we introduce the Length Value Model (LenVM), a token-level framework that models the remaining generation length at each decoding step. By formulating length modeling as a value estimation problem and assigning a constant negative reward to each generated token, LenVM predicts a bounded, discounted return that serves as a proxy for the remaining generation horizon. This formulation enables scalable, annotation-free value pretraining, as dense supervision signals can be automatically derived from sampled trajectories without human labeling. Extensive experiments demonstrate that LenVM provides a highly effective signal for inference time length control. On the LIFEBench exact length matching task, applying LenVM to a 7B model improves the length score from 30.9 to 64.8, significantly outperforming frontier closed-source models. Furthermore, LenVM enables continuous control over the trade off between performance and efficiency. On GSM8K at a budget of 200 tokens, LenVM maintains 63 percent accuracy compared to 6 percent for token budget baseline. It also accurately predicts total generation length from the prompt boundary. Finally, we show that the token level values of LenVM offer an interpretable lens into generation dynamics, revealing how specific tokens shifts between shorter and longer reasoning regimes.