Routing Entropy: A Hidden Self-Verifier for Free in Mixture-of-Experts LLMs
Abstract
Modern LLMs are observed to be overconfident in their own outputs, making them poor self-verifiers. We found that for mixture-of-experts (MoE) LLMs, the Routing Entropy (RE) is a more reliable self-verifier than output confidence (OC). In other words, more confident expert choices in MoE usually lead to higher-quality outputs. We examined RE for self-verification on multiple recent MoE LLMs. While RE consistently outperforms OC, combining them leads to the best performance. Moreover, our RE-based analysis shows that better MoEs are usually more confident in expert choices. These novel insights further motivate an RE-regularizer for post-training of MoE routers. Our extensive experiments demonstrate the superiority of RE for self-verification across diverse reasoning and knowledge benchmarks. Furthermore, RE-regularization in post-training significantly enhances MoE’s internal calibration and self-awareness, yielding a more reliable model without sacrificing its core generative performance.