Summary explanation: SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
10 Aug 2026 | Paper Review Code Refactoring AI Coding Agents BenchmarkContents
- Summary
- 1 Introduction
- 2 Related work
- 3 SWE-Bench ProMax
- 4 Experiments
- 5 Results and analysis
- 6 Conclusion
- Appendix
- Brief Thoughts
This article explains the key points of SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring.
- 2026-08-10 (arXiv), COLM 2026
- Shi, Yuling, Xu, Jinghan, Fu, Kelin, Zeng, Wenhao, He, Shilin, Zhang, Lei, Liu, Yue, Zhao, Zelin, Zhuo, Terry Yue, Cao, Jialun, et al.
- Shanghai Jiao Tong University, Peking University, The Hong Kong University of Science and Technology, Douyin Group, University of Chinese Academy of Sciences, National University of Singapore, Monash University
- Paper
- Project Page
Summary
- SWE-Bench ProMax is an expert-curated benchmark of 170 large-scale code-refactoring tasks from real commits in 70 repositories and seven languages: Python, Java, TypeScript, Go, C, C++, and Rust. It evaluates behavior-preserving, cross-file transformations by whether an agent’s final repository state passes the complete test suite; the best reported resolve rate is 41.2%.
1 Introduction
The paper argues that established coding benchmarks are becoming saturated and may have evaluation problems when tests are overly narrow or overly broad, task specifications are ambiguous, or public-repository tasks are contaminated. It presents refactoring as a long-horizon setting that requires coordinated changes across files while preserving behavioral invariants. For long-horizon setting, see this post.
The comparison quantifies the intended scale gap: 30% of ProMax instances modify more than 10 files and 32% change more than 200 LOC, versus 86% single-file and 93% at-most-20-LOC instances in SWE-bench Verified.
The NASA F'Prime example makes the cross-file requirement concrete: its refactor spans 244 files across templates, drivers, OS components, services, and build-related code rather than a localized edit.
2 Related work
Related work covers function-level generation, repository-level issue resolution, multilingual evaluation, long-horizon software tasks, and benchmark curation. The authors position ProMax as combining repository-level execution-based evaluation, multilingual refactoring tasks, large patches, and expert review of specifications and tests.
2.1 Coding benchmarks
The paper contrasts early function-level benchmarks, including HumanEval and MBPP, with repository-level suites such as SWE-bench, Multi-SWE-bench, SWE-bench Pro, SWE-EVO, and Terminal-Bench. It cites frontier scores above 75% on SWE-bench Verified and lower success on long-horizon tasks as motivation for a more demanding evaluation setting.
2.2 Code refactoring
The two refactoring-specific benchmarks discussed are limited either to Python or Java, smaller average changes, or the absence of human verification of test quality. Traditional refactoring tools mainly detect historical transformations, while LLM-refactoring studies use disparate measures such as code-smell counts, compilability, or alignment scores.
3 SWE-Bench ProMax
SWE-Bench ProMax is a multilingual, execution-based benchmark for measuring large behavior-preserving refactors. Each task combines a pre-refactoring repository state, a rewritten issue description, curated tests, and the original developer’s gold patch.
3.1 Task formulation
Each instance contains a pre-configured Docker environment, a precise natural-language specification, a test suite, and a gold patch. An agent receives the environment and specification, and an instance is resolved if and only if the agent’s final modifications pass every test; intermediate commands are not evaluated.
3.2 Dataset construction and curation
The three-stage pipeline selects GitHub repositories with at least 500 stars, an approved license, and a target language comprising at least 80% of the codebase; it then extracts post-January-2025 commits containing “refactor,” excluding “bug fix,” that modify both test and non-test files. After Docker validation, experts remove simple or insufficiently cross-file tasks, review inappropriate tests, rewrite specifications with LLM assistance, and perform final consistency checks; 170 of 29,782 candidates remain.
The matrix shows ProMax is the only listed benchmark marked simultaneously as execution-based, repository-level, multilingual, refactoring-focused, averaging more than five modified files, and expert-curated.
The pipeline visualizes the selection funnel from repository search and commit extraction through Docker validation to expert filtering, problem rewriting, and verification. It explains why the released set is much smaller than the initial candidate pool.
3.3 Dataset composition
The 170 tasks average 11.4 source-patch files, 261.6 source LOC, and 8,179.5 source-patch tokens; test patches add 4.5 files and 185.5 LOC on average. The total average is 15.9 modified files, issue descriptions average 685.3 tokens, and the largest source patch modifies 182 files.
4 Experiments
Experiments compare six proprietary and open-weight models under mini-swe-agent and OpenHands. Evaluation uses all 170 instances, pre-built isolated Docker environments, a 300-step limit, a $10 per-instance cost limit, Pass@1 resolve rate, per-language resolve rates, agent steps, and average API cost.
4.1 Experimental setup
mini-swe-agent provides iterative file viewing, editing, search, and bash execution, whereas OpenHands provides sandboxed command execution and structured file-editing tools. Using the same scaffolds and limits across models is intended to support a controlled comparison of model and runtime-tooling effects.
4.2 Models
The proprietary models are Gemini-3-Pro, Claude Sonnet 4.6, and GPT-5.2; the open-weight models are GLM-5, Kimi-K2.5, and Qwen3.5. The selection spans model families, architectures, and training paradigms for a common large-scale refactoring workload.
5 Results and analysis
ProMax remains difficult under the reported setups: GPT-5.2 with OpenHands attains the highest overall resolve rate, 41.2%, below the 75%+ frontier performance cited for SWE-bench Verified. Results vary substantially with agent scaffold, programming language, and API cost, and no model leads in every language.
5.1 Main results
Under OpenHands, GPT-5.2 reaches 41.2%, Claude Sonnet 4.6 reaches 38.8%, and GLM-5 and Qwen3.5 each reach 36.5%; GPT-5.2 rises from 21.8% with mini-swe-agent. Language leaders are GPT-5.2 for Python (48.3%) and C (75.0%), Claude Sonnet 4.6 for TypeScript (53.6%) and Rust (63.6%), GLM-5 for Java (34.6%), Kimi-K2.5 for Go (43.5%), and Qwen3.5 for C++ (54.5%).
The table supplies the numerical basis for claims about scaffold sensitivity, language variation, and cost efficiency. It reports OpenHands GPT-5.2 at 41.2% overall and OpenHands GLM-5 at 36.5% for $0.24 per instance.
5.2 Agent behavior analysis
For Claude Sonnet 4.6 and Kimi-K2.5, agent file-modification distributions track gold patches on small tasks but under-cover larger patches: agents reach 90% cumulative coverage at about 10 modified files, whereas gold patches reach it at about 20. Failed runs also use more interaction rounds, consistent with incomplete cross-file propagation and unproductive edit–revert cycles.
The left cumulative distributions show that Claude Sonnet 4.6 and Kimi-K2.5 modify fewer files than the gold patches, especially for larger changes. The right panel separates passing and failing runs, showing that failures tend to require more interaction rounds.
5.3 Cost and efficiency analysis
Cost is not proportional to resolve rate: under OpenHands, Claude Sonnet 4.6 costs $4.77 per instance for 38.8%, GLM-5 costs $0.24 for 36.5%, and GPT-5.2 reaches 41.2% at $3.60. Qwen3.5 takes the most average steps under both scaffolds, 155.4 and 141.2, without leading overall; the paper also states that repositories use approved open-source licenses and discloses LLM assistance for curation, appendix classification, and paper editing.
6 Conclusion
The conclusion attributes the benchmark’s difficulty to its large cross-file scope and quality-controlled specifications and tests. It identifies incomplete propagation of changes across affected files as a central empirical limitation of current agents and reports that open-weight models can approach leading proprietary performance at much lower API cost.
The repository inventory documents the public code sources and licenses used in the benchmark, supporting the stated breadth of 70 repositories across seven languages.
Per-language statistics reveal uneven patch complexity: C++ and Java average 21.4 and 20.8 modified files, while C has the highest average LOC at 424.1. TypeScript has 28 instances but draws from only two repositories, a concentration caveat identified by the paper.
Every task has at least two category labels, and 46.5% have three or more, showing that the benchmark contains compound software-maintenance changes rather than isolated transformations.
The selected examples show the benchmark's range: the listed C++ instance changes 244 files, while the TypeScript example changes 27 files. Their summaries cover header unification, temporal resolution, configuration renaming, tensor APIs, stream buffering, provider formatting, and component APIs.
The co-occurrence matrix shows that API Interface Change and Refactoring Cleanup overlap in 79 instances. Bug Fix also co-occurs frequently with Refactoring Cleanup (48) and API Interface Change (33), showing that tasks commonly combine several maintenance concerns.
The skill labels identify cross-file reasoning in 169 instances and API semantics in 168, followed by interface-contract reasoning in 165 and pattern matching in 156. These annotations support the benchmark's emphasis on repository-wide coordination.
Appendix
- The appendix lists the 70 source repositories and licenses, gives per-language patch statistics, and analyzes task categories and required skills. Refactoring Cleanup and API Interface Change occur in 66.5% and 65.3% of instances, while cross-file reasoning and API semantics are labeled for 99.4% and 98.8%; representative tasks range from a 27-file TypeScript API unification to a 244-file C++ header refactor.
Brief Thoughts
The benchmark’s contribution is the attempt to align task specifications, test suites, and behavior-preserving outcomes for large, multilingual refactoring tasks. Its 41.2% best result is conditional on the evaluated scaffolds and budgets, so it should be read as a benchmark outcome rather than a direct measure of all software-engineering capability; ongoing auditing of curation decisions, repository concentration, and test sufficiency remains important.