AgentIR: Reasoning-Aware Retrieval for Deep Research Agents
Abstract
Deep Research agents are rapidly emerging as primary users of modern retrieval systems. However, existing retrievers remain designed for standalone queries from human users or single-turn Retrieval-Augmented Generation. In contrast, Deep Research is a multi-turn reasoning process, where retrieval should condition on the agent’s evolving reasoning state, rather than the query alone. To address this, we introduce: (1) Reasoning-Aware Retrieval, a retrieval paradigm that jointly embeds the agent's reasoning trace alongside its query; and (2) DR-Synth, a data synthesis method that generates Deep Research retriever training data from standard QA datasets. We demonstrate that both components are independently effective, and their combination yields a trained embedding model, AgentIR-4B, with substantial gains. On the challenging BrowseComp-Plus benchmark, AgentIR-4B achieves 68% accuracy with the open-weight agent Tongyi-DeepResearch, compared to 52% with conventional embedding models twice its size, and 37% with BM25. Code and data will be released.