Distractor Injection Attacks on Large Reasoning Models: Characterization and Defense
Abstract
Recent advances in large reasoning models (LRMs) have enabled remarkable performance by generating long Chain-of-Thought (CoT) reasoning traces. In this paper, we identify and systematically analyze a critical vulnerability we term reasoning distraction, where LRMs are diverted from their primary objective by irrelevant yet complex tasks embedded in the prompt. While such attacks can be instantiated via prompt injection, we show that they expose a distinct reasoning-level failure mode that is not captured by existing work, which primarily focuses on output manipulation rather than corruption of intermediate reasoning. Reasoning distraction can induce covert compliance, where the model silently follows adversarial objectives within its hidden reasoning trace while maintaining well-formed, seemingly correct outputs. This poses a significant risk to automated agentic pipelines that rely on final outputs without access to internal reasoning. We introduce DistractBench, a comprehensive benchmark for evaluating robustness to reasoning distraction, and show that even state-of-the-art LRMs are highly vulnerable, with accuracy drops of up to 60%. To mitigate this vulnerability, we propose a training-based defense using Supervised Fine-Tuning (SFT) and preference optimization on Distract-SFT and Distract-DPO, two synthetically generated datasets designed to improve robustness to reasoning-level corruption, achieving gains of over 50 points. Together, these components form the Distraction Suite, a standardized toolkit for evaluating and improving reasoning robustness in LRMs.