From SWE-ZERO to SWE-HERO: Execution-free to Execution-based Fine-tuning for Software Engineering Agents
Abstract
To address the complexities of automated software issue resolution, we present SWE-Zero to SWE-Hero, a high-performance SFT recipe that achieves state-of-the-art resolution rates on SWE-bench. Traditional methods often depend on verifiable execution environments, which are labor-intensive and technically complex to scale. In contrast, our approach pushes the boundaries of SFT by distilling open-weight frontier LLMs through a streamlined two-stage pipeline: (1) SWE-Zero, which leverages execution-free fine-tuning at scale, and (2) SWE-Hero, which utilizes lightweight, execution-based SFT to refine agent performance. Our empirical results establish a new state-of-the-art among open-source models of comparable size on SWE-bench Verified. We release a comprehensive dataset of 300k \emph{zero} and 13k \emph{hero} agent trajectories distilled from Qwen3-Coder-480B, alongside a suite of agents fine-tuned from the Qwen2.5-Coder series. Notably, SWE-Zero-32B achieves a 62.2\% resolution rate. We further validate the paradigm’s generalizability on SWE-bench Multilingual, where SWE-Hero-32B reaches 44.1\%. This strong zero-shot performance across languages, achieved despite Python-only training, underscores the robustness of our distillation recipe.