Reasoning Fine-Tuning Induces Persistent Latent Policy States
Abstract
Reasoning-specialized language models show large performance gains over base models, yet the internal changes responsible for improved multi-step reasoning remain poorly understood. It is unclear whether reasoning fine-tuning primarily improves local token-level competence or induces global changes in how models organize inference over time. We address this by modeling Chain-of-Thought reasoning as a switching dynamical system (SDS), in which internal representations evolve under discrete, persistent latent policy states. We introduce a framework combining time-aware contrastive representation learning with discrete regime discovery to recover latent policies directly from activation trajectories. Across multiple reasoning benchmarks and model scales, we find a clear distinction between base and fine-tuned models: base models exhibit weaker, less consistent regime structure with lower persistence and reduced state utilization, while reasoning-fine-tuned models reliably transition between a small number of persistent states with structured dynamics. These regimes exhibit functional specialization aligned with reasoning stages including planning, computation, verification, and coordination. We further identify a finite internal policy capacity beyond which persistence and predictive power degrade. Causal interventions show that latent regimes are actionable, as ablating specific regimes disrupts reasoning, while steering base models with the reasoning model's policy SDS improves performance. Our results suggest that reasoning fine-tuning induces a global reorganization of internal dynamics, giving rise to a latent reasoning policy governing sustained inference over extended horizons and offering a new lens for mechanistic interpretability of reasoning models.