Variational Co-Evolution via Reinforcement Learning
Abstract
Despite recent progress, large language models still exhibit limitations and remain brittle on challenging reasoning tasks, as effective reasoning depends not only on the underlying solution policy but also on the capability to construct prompts that reliably elicit correct reasoning. Most existing methods optimize these two components separately, focusing either on prompt rewriting or solely on strengthening the solving policy, which makes it hard to build a persistent, mutually reinforcing loop between prompting and solution generation. To address this limitation, we propose Variational Co-Evolution (VCE), a novel co-evolving framework that jointly improves prompt generation and solution generation with the explicit goal of strengthening both capabilities simultaneously. VCE formulates co-evolution through a reinforcement learning framework with a variational inference, treating prompts as adaptive decisions jointly optimized with solution policies. By enabling coherent updates of prompting and reasoning under a unified objective, VCE provides a stable and scalable paradigm for co-evolving prompts and solvers on challenging reasoning tasks. Experiments across multiple benchmarks demonstrate strong and consistent gains, validating the effectiveness of co-evolution for improving both prompt construction and solution accuracy.