Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration
Abstract
Preference alignment often makes large language models (LLMs) overconfident and poorly calibrated. Traditional post-hoc temperature scaling is inherently dataset-dependent: a temperature fitted on one validation set does not generalize across domains. This motivates us to modify model parameters during training to improve calibration. We propose maximizing the entropy of predictive distributions as the calibration objective, which directly targets overconfidence by discouraging overly concentrated predictions. Inspired by temperature scaling, we realize this through a bilevel optimization formulation, where the lower level trains the model under a parametric loss and the upper level selects loss hyperparameters on held-out inputs to maximize entropy. To make the framework practical at LLM scale, we adopt an efficient first-order approximation that avoids explicit second-order computation. Experiments demonstrate that our method yields well-calibrated LLMs with particular advantages in out-of-domain generalization.