Introspective Diffusion Language Models
Abstract
Diffusion language models (DLMs) offer a compelling promise: parallel token generation could break the sequential bottleneck of autoregressive (AR) decoding. Yet in practice, DLMs consistently lag behind AR models in quality. We argue that this gap stems from a fundamental failure of introspective consistency: AR models agree with what they generate, whereas DLMs often do not. We formalize this via the introspective acceptance rate, which quantifies whether a model internally accepts its previously generated tokens. Through this lens, we uncover a key structural advantage of AR models: causal masking combined with logit shifting implicitly enforces introspective consistency during training. Motivated by this insight, we introduce the Introspective Diffusion Language Model (I-DLM), a new paradigm that preserves the introspective consistency of AR training while retaining the diffusion-style parallelism. I-DLM uses a novel introspective strided decoding (ISD) algorithm, which enables the model to verify previously generated tokens while advancing new ones in the same forward pass. This yields a new quality–efficiency frontier unavailable to either AR or prior diffusion models, with stride providing a controllable tradeoff between verification depth and parallel progress. Empirically, I-DLM-8B is the first DLM to match the quality of its same-scale AR counterpart while surpassing all prior DLMs in both quality and practical serving efficiency across 15 benchmarks. It attains 72.5 on AIME-24 and 45.1 on LiveCodeBench-v6, outperforming LLaDA-2.1-mini (16B) by more than 29 and 14 points, respectively. Finally, with a single-pass self-speculative decoding pipeline and gated LoRA, ISD enables lossless acceleration and better efficiency at high concurrency compared to speculative decoding.