CONCORD: Label-Free Calibration of Verbalized LLM Confidence via Rollout Consistency
Abstract
Calibrated confidence, where a model's stated probability matches its actual expected accuracy, is often essential in contexts where confidence scores are being used by human users to make decisions or are provided as inputs to downstream agents. Existing approaches for training models to verbalize calibrated confidence require ground-truth labels, limiting their applicability under distribution shift. Multi-pass consistency methods achieve good calibration but significantly increase inference costs. We introduce CONCORD, to our knowledge the first label-free method for training language models to generate calibrated verbalized confidence scores. Adapting models with Group Relative Policy Optimization already generates multiple rollouts per prompt; we repurpose these rollouts as a calibration signal by deriving confidence targets from semantic agreement among sampled answers, using either Natural Language Inference-based answer clustering or semantic entropy. The resulting reward encourages the model to align its stated confidence with empirical semantic agreement, without requiring ground-truth labels at any stage of training. At inference, the model generates a reasoning trace, an answer, and a calibrated confidence score. Across 4 benchmarks and four models (Qwen-2.5 and Llama-3, 3B--8B), CONCORD achieves the lowest expected calibration error among verbalized confidence methods on most model–task pairs (e.g., 0.075 vs. 0.399 on GSM8K), while maintaining competitive AUROC performance.