MELD: Multilingual Ensemble via Logical Debate
Abstract
Large language models can benefit from inference-time strategies such as self-consistency and ensemble-based methods, including multi-agent debate. Recent work, however, suggests that the effectiveness of these approaches depends on the diversity of the reasoning trajectories they generate. In this work, we consider multilinguality as a practical way to induce such diversity. Whereas monolingual reasoning paths often remain highly correlated, reasoning across multiple languages can encourage exploration of a broader solution space. Motivated by this idea, we propose MELD, Multilingual Ensemble via Logical multi-agent Debate, which leverages language variation to produce complementary reasoning trajectories. Across multiple benchmarks and model series, MELD achieves competitive and consistently favorable performance compared with baselines, and our analyses suggest that these gains are associated with improved error complementarity, aggregation, and debate-based correction under multilingual reasoning.