Lost in Distraction: LLM Planning Agents Recognize but Fail to Utilize Relevant Information under Context Noise
Abstract
Large Language Model (LLM)-based planning agents operate under long, dynamically accumulated contexts that inevitably introduce noise. While it is commonly assumed that such noise impairs reasoning, our study shows that its primary impact lies elsewhere. Through a controlled noise injection framework on the DeepPlanning benchmark, we observe that noise does not significantly degrade the model’s ability to identify relevant information. Instead, it leads to systematic failures in utilizing that information, resulting in plans that are factually correct but violate specific user requirements. Noise also alters agent behavior, leading to reduced tool usage and fewer decision steps. To explain this discrepancy, we propose a two-stage framework for agent behavior under noise: information recognition and information utilization. Empirically, we identify distinct failure modes: some models fail to reliably identify relevant information, while others successfully identify it but fail to consistently apply it in decision-making. Furthermore, we find that selection mechanisms improve performance only when recognition is reliable, and can even be detrimental otherwise. These findings reveal a fundamental gap between recognizing and utilizing relevant information, highlighting consistent information utilization as a key bottleneck for developing robust LLM agents.