SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
Abstract
Code refactoring is among the most frequent and consequential activities in professional software development, yet existing benchmarks overlook it in favor of bug fixing and feature implementation, and suffer from critical issues including data contamination, overly narrow and overly broad tests, and limited language coverage. We introduce SWE-Cascade, a challenging, long-horizon benchmark of 196 expert-curated code refactoring instances drawn from real commits in actively maintained GitHub repositories across seven programming languages (Python, Java, TypeScript, Go, C, C++, and Rust). Every instance undergoes rigorous, multi-stage curation: issue descriptions are rewritten from scratch to provide precise, unambiguous specifications, and test suites are manually reviewed to eliminate tests that are overly narrow (rejecting valid solutions) or overly broad (checking unstated requirements). The resulting instances average 11.8 modified files and 355.5 lines of code, substantially exceeding the scale of existing benchmarks. Evaluation with frontier models shows that the best-performing model achieves only 28.9% resolve rate, confirming that SWE-Cascade presents a meaningful and unsaturated challenge for current AI coding agents.