Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts
Ellis Brown ⋅ Jihan Yang ⋅ Shusheng Yang ⋅ Rob Fergus ⋅ Saining Xie
Abstract
Robust benchmarks are crucial for accurately evaluating Multimodal Large Language Models (MLLMs). However, we find that models can ace many multimodal benchmarks *without* strong visual understanding by exploiting biases, linguistic priors, and superficial patterns. This is particularly problematic for *vision-centric* benchmarks, which explicitly aim to require visual inputs to be solved. We introduce a diagnostic principle for robust benchmark design: if a benchmark *can* be gamed, it *will* be. Therefore, designers should proactively try to "game" their own benchmarks first as a key step in the development lifecycle---adopting rigorous diagnostic and debiasing procedures to systematically identify, quantify, and mitigate non-visual biases. We demonstrate that effective diagnosis of these issues *must* involve directly "training on the test set"---i.e., probing the *specific test set* being released for its intrinsic, exploitable patterns. To demonstrate an effective realization of this standard, we propose a systematic approach involving two core components: First, we *diagnose* benchmark susceptibility using a "Test-set Stress-Test" (TsT) methodology. The primary diagnostic tool involves fine-tuning a powerful Large Language Model (LLM) via $k$-fold cross-validation on *exclusively* the non-visual, textual inputs of the test set to unveil shortcut performance and derive a quantitative, sample-level bias score, $s(x)$. We complement this with a lightweight Random Forest-based diagnostic trained on hand-crafted features, enabling rapid auditing and interpretable bias analysis. Second, we *debias* benchmarks by systematically filtering samples identified as highly biased according to $s(x)$ using an "Iterative Bias Pruning" (IBP) procedure. Applying this framework to four prominent benchmarks---VSI-Bench, CV-Bench, MMMU, and VideoMME---we reveal that non-visual shortcuts are both pervasive and heterogeneous: on template-based benchmarks, TsT uncovers dramatic learnable patterns (up to +33 points from text-only fine-tuning) that exceed even GPT-4o in blind evaluation; on knowledge-intensive MMMU, TsT correctly identifies near-zero learnable shortcuts, while frontier models achieve over 50% blind accuracy through world knowledge alone. We apply our full debiasing framework to both VSI-Bench (template-based) and MMMU (non-template, knowledge-intensive), demonstrating that IBP generalizes across benchmark types: on VSI-Bench-Debiased, the vision-blind gap widens significantly, while on MMMU, mean bias scores decrease by 49% with only 11% sample removal.
Successful Page Load