The Depth Ceiling: On the Limits of Large Language Models in Discovering Latent Planning
Abstract
The viability of chain-of-thought (CoT) monitoring crucially hinges on models being unable to reason effectively in their latent representations. Yet little is known about the limits of such latent reasoning in LLMs. We test these limits by studying whether models can discover multi-step planning strategies without supervision on intermediate steps and execute them latently within a single forward pass. Using graph path-finding tasks that precisely control the number of required latent planning steps, we uncover a striking limitation unresolved by massive scaling: tiny transformers trained from scratch discover up to three latent steps, a fine-tuned GPT-4o reaches five, and GPT-5.4 attains seven under few-shot prompting. Yet a model fine-tuned only on five-step problems discovers a solution that generalizes to up to eight steps at test-time, demonstrating a dissociation between the ability to discover and execute latent planning strategies. Notably, the task is trivial when models are allowed CoT. If similar limits hold more broadly, strategies requiring multiple coordinated latent planning steps may need to be explicitly taught or externalized, lending credence to CoT monitoring.