Cross-Model Disagreement as a Label-Free Correctness Signal
Abstract
Detecting when a language model is wrong without ground truth labels is a fundamental challenge for safe deployment. Existing approaches rely on a model's own uncertainty, such as token entropy or confidence scores, but these signals fail critically on the most dangerous failure mode: confident errors, where a model is wrong but certain. In this work we introduce \textit{cross-model disagreement} as a correctness indicator — a simple, training-free signal that can be dropped into existing production systems, routing pipelines, and deployment monitoring infrastructure without modification. Given a model's generated answer, cross-model disagreement computes how surprised or uncertain a second verifier model is when reading that answer via a single forward pass. No generation from the verifying model is required, and no correctness labels are needed. We instantiate this principle as \gls{cmp}, which measures the verifying model's surprise at the generating model's answer tokens, and \gls{cme}, which measures the verifying model's uncertainty at those positions. Both \gls{cmp} and \gls{cme} outperform within-model uncertainty baselines across benchmarks spanning reasoning, retrieval, and mathematical problem solving (MMLU, TriviaQA, and GSM8K). On MMLU, \gls{cmp} achieves a mean AUROC of 0.75 against a within-model entropy baseline of 0.59. These results establish cross-model disagreement as a practical, training-free approach to label-free correctness estimation, with direct applications in deployment monitoring, model routing, selective prediction, data filtering, and scalable oversight of production language model systems. Code is provided in the supplementary materials and will be publicly released upon acceptance.