Measuring Nines of Reliability: Sample-Efficient LLM Evaluation in Saturated GSM Benchmarks
Eungyeup Kim ⋅ Chenchen Gu ⋅ Vashisth Tiwari ⋅ J Kolter
Abstract
While existing benchmarks demonstrate the near-perfect performance of large language models (LLMs) on various tasks, this apparent saturation obscures the need for rigorous evaluation of their reliability. In real-world deployment, distinguishing between "five-nines" (99.999%) and "three-nines" (99.9%) reliability is critical, since these correspond to an order-of-magnitude gap in error rate. However, standard Monte Carlo evaluation under uniform sampling requires prohibitively large sample sizes to achieve tight confidence intervals around rare failure probabilities, making evaluation computationally expensive. In this paper, we show that LLM failures exhibit strong systematic patterns: across broad parameterized input spaces, a small subset of inputs disproportionately accounts for the majority of failures. Leveraging this observation, we apply the cross-entropy (CE) method to learn a sampling distribution concentrated on failure-prone inputs, then use importance sampling to obtain unbiased failure-rate estimates with tight confidence intervals using significantly fewer inferences. On parameterized GSM8K templates across reasoning LLMs with self-consistency decoding, our approach achieves up to an $19.19\times$ reduction in inferences for equivalent confidence bounds compared to naive uniform sampling. Furthermore, failure patterns transfer across different test-time compute scales (e.g., varying numbers of generations for majority voting), enabling efficient estimation for high-compute models using sampling distributions learned in low-compute settings. By enabling evaluation of rare failure rates with orders-of-magnitude fewer samples, our framework provides a critical tool for ensuring that frontier LLM deployments meet the reliability standards required for high-stakes applications.
Successful Page Load