ComDoc: A Comprehensive Benchmark for Document Retrieval-Augmented Generation
Abstract
Document Retrieval-Augmented Generation (DocRAG) is an emerging paradigm crucial for complex tasks such as financial report analysis and multi-source knowledge retrieval. However, existing benchmarks typically target isolated capabilities, such as multi-modal or multi-hop reasoning, and offer only single-perspective evaluations, revealing a need for more comprehensive assessment. In this paper, we present ComDoc, a large-scale multi-modal DocRAG benchmark spanning single-hop, multi-hop, and multi-page tasks. ComDoc includes 1,755 real-world documents from four domains, comprising 103,877 pages and 5,561 QA instances. To construct ComDoc, we design a three-stage pipeline: data collection, hierarchical QA pairs generation, and quality control. Using ComDoc, we systematically evaluate state-of-the-art encoder-based, LLM-based, and MLLM-based retrieval models, as well as the generative capabilities of leading open-source LLMs and MLLMs. Our results demonstrate substantial performance gaps across different models and task types in both retrieval and generation. We will release ComDoc and the evaluation code to facilitate further research.