Message Passing Enables Efficient Reasoning
Xuecheng Liu ⋅ Daman Arora ⋅ Gokul Swamy ⋅ Andrea Zanette
Abstract
While inference-time scaling has improved the reasoning abilities of large language models (LLMs), the need to generate long chains-of-thought (CoTs) is a computational bottleneck. Thus, in contrast to sequential scaling methods like CoT, recent parallel scaling techniques instead use fork and join (FJ) primitives to divide work across multiple LLM threads. Although conceptually promising, this necessitates a centralized controller, limiting scalability. We introduce Message Passing Language Models (MPLMs), a framework for LLM reasoning in which threads communicate directly via lightweight send and receive primitives. MPLMs enable efficient scaling through two key mechanisms: (1) reduced communication costs, achieved by avoiding redundant context sharing, and (2) preemption, which allows threads to terminate early based on partial information from their peers. In Sudoku puzzles, we show that MPLMs require an asymptotically smaller context than both serial CoT and parallel FJ . Most importantly, we fine-tune a single model to solve 25 $\times$ 25 puzzles that remain challenging for standard CoT and FJ approaches, as well as frontier reasoning models without tools. We also demonstrate that on 3-SAT puzzles, where search dominates, the capability of preemption allows termination of unpromising branches, which results in improved efficiency. Together, these results demonstrate that MPLMs could provide a principled and scalable alternative to existing sequential and centrally coordinated parallel reasoning paradigms.
Successful Page Load