WebReal: Benchmarking DeepSearch Agents on Real-world User Information Needs
Abstract
Evaluating web browsing agents demands benchmarks that faithfully reflect how real users seek information online. We argue that existing benchmarks suffer from a \textbf{structural misalignment} with genuine user needs: their reverse-engineered, puzzle-style queries exercise a narrow slice of web-searching capabilities while leaving the majority of authentic information needs untested. To characterize this gap, we develop an \textbf{operation-oriented taxonomy} comprising six cognitive operations: \textit{direct & Implicit retrieval}, \textit{chain traversal}, \textit{document extraction}, \textit{computation}, \textit{disambiguation}, and \textit{temporal linking}. Applying this taxonomy reveals severe distributional skew in existing benchmarks, over \textbf{64\%} of BrowseComp queries reduce to disambiguation alone. To close this gap, we introduce \textbf{WebReal}, a benchmark of \textbf{507} human-verified QA pairs derived from authentic user requests, spanning \textbf{15} domains across all six operation types. Experiments across state-of-the-art agents reveal that WebReal exposes capability profiles and ranking shifts invisible to existing benchmarks: models with near-identical aggregate scores exhibit divergent strengths across operation types, and performance degrades sharply with reasoning complexity in operation-type-specific patterns. These findings demonstrate that WebReal enables a finer-grained decomposition of web search capability than any prior benchmark. Our benchmark will be publicly available soon.