ResearcherBench: Evaluating Deep AI Research Systems on Open-ended AI Research Tasks
Abstract
Deep research agents have emerged with capabilities extending from basic queries to sophisticated research tasks. However, existing benchmarks primarily evaluate their abilities on web retrieval and report generation, overlooking the potential for discovering, integrating and generating insights in AI research. To address this gap, we introduce ResearcherBench, the first benchmark focused on evaluating the capabilities of Deep AI Research Systems (DARS) on open-ended AI research tasks. We curated a dataset of 65 questions expertly selected from real-world AI research scenarios, spanning 35 different AI subjects and categorized into 3 different types. Our dual evaluation framework consists of two components: rubric assessment which implements expert-designed rubric to evaluate insight quality, and factual assessment measuring citation accuracy (faithfulness) and coverage (groundedness). We evaluated several leading commercial DARS and baseline systems. Our evaluation results reveal that these systems excel at open-ended consulting questions compared to technical implementation questions, demonstrating the potential for DARS to serve as AI research partners. We open-source ResearcherBench and evaluation framework to accelerate the development of AI research assistants capable of scientific collaboration.