FERA: Uncertainty-Aware Federated Reasoning for Large Language Models
Abstract
Large Language Models (LLMs) exhibit strong multi-step reasoning, but their inference-time performance depends critically on the quality of available reasoning traces. In many realistic settings, such traces are generated by multiple clients holding private and heterogeneous data, motivating the problem of federated reasoning, where a server aggregates client-generated reasoning without centralized training or data sharing. A key challenge is that client reasoning quality is highly query-dependent, and naive aggregation can amplify miscalibrated confidence and correlated reasoning errors. We propose Uncertainty-Aware Federated Reasoning (FERA), a training-free framework that aggregates client reasoning using query-dependent trust weights derived from uncertainty signals, rather than treating all client outputs equally. We instantiate FERA with Uncertainty-Aware Self-Critique Aggregation (UA-SCA), which leverages structured self-critique and cross-client verification to identify unsupported reasoning steps and downweight high-confidence but weakly justified traces. We provide theoretical guarantees on the convergence of the resulting trust-weighted aggregation and demonstrate on multiple reasoning benchmarks that FERA consistently improves accuracy over standard ensembling baselines, while requiring minimal communication and no model updates.