Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
Abstract
Many important applications of long-context models require reasoning over large collections of items—summarizing trends across thousands of customer interactions, identifying patterns in research corpora, or analyzing behavioral data at scale. Evaluating whether models can analyze items atomically and aggregate findings requires information-dense tasks that make use of the full context, like writing a summary or describing trends across a collection. Yet these types of tasks often result in long-form outputs which are challenging to evaluate; as a result, current long-context benchmarks predominantly test retrieval-style tasks, where the correct answer depends on a small fraction of tokens and the rest is noise. We introduce OOLONG, a benchmark of distributional reasoning tasks that require analyzing individual chunks of text atomically and aggregating results. OOLONG includes synthetic tasks (OOLONG-synth) that enables controlled ablation of reasoning components, and a downstream setting (OOLONG-real) that requires reasoning over realistic conversational data. Tasks involve classification, counting, filtering by metadata, and temporal reasoning. Even frontier models struggle at this task at moderate context lengths, suggesting OOLONG captures capabilities not yet addressed by scale alone. We release the data and evaluation harness for OOLONG to enable further development of models that can reason over large quantities of text.