Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees
Sher Badshah ⋅ Ali Emami ⋅ Hassan Sajjad
Abstract
LLM-as-a-judge has become a widely adopted paradigm for scalable evaluation of language model outputs. However, when applied to objective, factual question answering in a reference-free setting, LLM judges must rely entirely on their parametric knowledge to assess correctness, making them vulnerable to hallucination and knowledge gaps that produce silent evaluation errors. While uncertainty quantification methods can flag unreliable verdicts, heuristic thresholds offer no formal guarantee on the error rate among accepted evaluations. We propose a risk-controlled selective evaluation framework that calibrates statistically valid uncertainty thresholds on a held-out set, ensuring that the false discovery rate (FDR) among non-abstained verdicts is bounded by a user-specified level~$\alpha$ with high probability. Rather than defaulting to abstention when uncertain, the framework routes uncertain instances to a retrieval-augmented mode: the judge searches the web for relevant evidence and re-evaluates with the retrieved context. A second calibrated threshold governs this mode, and we show that the Clopper--Pearson finite-sample FDR guarantee extends to the joint two-threshold routing without additional assumptions. Experiments across open-domain QA benchmarks and judge models of varying scales demonstrate rigorous error rate control with substantially improved coverage over single-mode baselines, with retrieval triggered selectively only when the judge is not confident enough to evaluate on its own.
Successful Page Load