What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
Meera Desai ⋅ Sang Truong ⋅ Hanna Wallach ⋅ Alex Chouldechova ⋅ A. Feder Cooper ⋅ Jean Garcia-Gathright ⋅ Daniel Ho ⋅ Abigail Z Jacobs ⋅ Sanmi Koyejo ⋅ Nicholas Pangakis ⋅ Angelina Wang
Abstract
Evaluation benchmarks play a central role in the development and governance of AI systems, yet persistent concerns about val concerns about their validity remain. It is unclear that such benchmarks actually measure the concepts they claim to assess. Drawing inspiration from the principles of $\textit{convergent and discriminant validity}$, we analyze 56 safety and capability benchmarks across 53 models using correlation analyses and item-level prediction models. When studying benchmarks purporting to measure the same construct, we often see inconsistent consistent rankings on the models. This suggests that the constructs, such as safety detection and exaggerated refusal, may not be sufficiently well-defined. When comparing across constructs, we find that some concepts are distinguishable from others (e.g., unsafe generation is different from bias), but many concepts (e.g., reasoning, knowledge, and comprehension) are not well separated. In some cases, evaluation format and task structure drive correlations between benchmarks more than the concepts they aim to measure. Finally, we investigate individual benchmarks to identify benchmarks that likely measure a different concept than reported. In particular, we find that BBQ, one of the most widely used bias benchmarks, is more correlated with reasoning benchmarks than with other bias benchmarks. Overall, we encourage benchmark developers and practitioners to evaluate benchmarks through the lens of convergent and discriminant validity when developing and using benchmarks. We release our extensive item-level dataset of model predictions to support future empirical work on benchmark validity.
Successful Page Load