Gorio Tech Blog search

Summary explanation: SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring

|

Contents

This article explains the key points of SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring.

  • 2026-08-10 (arXiv), COLM 2026
  • Shi, Yuling, Xu, Jinghan, Fu, Kelin, Zeng, Wenhao, He, Shilin, Zhang, Lei, Liu, Yue, Zhao, Zelin, Zhuo, Terry Yue, Cao, Jialun, et al.
  • Shanghai Jiao Tong University, Peking University, The Hong Kong University of Science and Technology, Douyin Group, University of Chinese Academy of Sciences, National University of Singapore, Monash University
  • Paper
  • Project Page

Read this article in Korean


Summary

  • SWE-Bench ProMax is an expert-curated benchmark of 170 large-scale code-refactoring tasks from real commits in 70 repositories and seven languages: Python, Java, TypeScript, Go, C, C++, and Rust. It evaluates behavior-preserving, cross-file transformations by whether an agent’s final repository state passes the complete test suite; the best reported resolve rate is 41.2%.

1 Introduction

The paper argues that established coding benchmarks are becoming saturated and may have evaluation problems when tests are overly narrow or overly broad, task specifications are ambiguous, or public-repository tasks are contaminated. It presents refactoring as a long-horizon setting that requires coordinated changes across files while preserving behavioral invariants. For long-horizon setting, see this post.

The comparison quantifies the intended scale gap: 30% of ProMax instances modify more than 10 files and 32% change more than 200 LOC, versus 86% single-file and 93% at-most-20-LOC instances in SWE-bench Verified.

Stacked distributions of modified-file counts and modified LOC per instance for SWE-Bench ProMax, SWE-bench Pro, and SWE-bench Verified.
Stacked distributions of modified-file counts and modified LOC per instance for SWE-Bench ProMax, SWE-bench Pro, and SWE-bench Verified.

The NASA F'Prime example makes the cross-file requirement concrete: its refactor spans 244 files across templates, drivers, OS components, services, and build-related code rather than a localized edit.

Excerpted file list for a C++ NASA F'Prime refactoring instance with 244 modified files and +591 / −514 lines.
Excerpted file list for a C++ NASA F'Prime refactoring instance with 244 modified files and +591 / −514 lines.

Related work covers function-level generation, repository-level issue resolution, multilingual evaluation, long-horizon software tasks, and benchmark curation. The authors position ProMax as combining repository-level execution-based evaluation, multilingual refactoring tasks, large patches, and expert review of specifications and tests.

2.1 Coding benchmarks

The paper contrasts early function-level benchmarks, including HumanEval and MBPP, with repository-level suites such as SWE-bench, Multi-SWE-bench, SWE-bench Pro, SWE-EVO, and Terminal-Bench. It cites frontier scores above 75% on SWE-bench Verified and lower success on long-horizon tasks as motivation for a more demanding evaluation setting.

2.2 Code refactoring

The two refactoring-specific benchmarks discussed are limited either to Python or Java, smaller average changes, or the absence of human verification of test quality. Traditional refactoring tools mainly detect historical transformations, while LLM-refactoring studies use disparate measures such as code-smell counts, compilability, or alignment scores.

3 SWE-Bench ProMax

SWE-Bench ProMax is a multilingual, execution-based benchmark for measuring large behavior-preserving refactors. Each task combines a pre-refactoring repository state, a rewritten issue description, curated tests, and the original developer’s gold patch.

3.1 Task formulation

Each instance contains a pre-configured Docker environment, a precise natural-language specification, a test suite, and a gold patch. An agent receives the environment and specification, and an instance is resolved if and only if the agent’s final modifications pass every test; intermediate commands are not evaluated.

3.2 Dataset construction and curation

The three-stage pipeline selects GitHub repositories with at least 500 stars, an approved license, and a target language comprising at least 80% of the codebase; it then extracts post-January-2025 commits containing “refactor,” excluding “bug fix,” that modify both test and non-test files. After Docker validation, experts remove simple or insufficiently cross-file tasks, review inappropriate tests, rewrite specifications with LLM assistance, and perform final consistency checks; 170 of 29,782 candidates remain.

The matrix shows ProMax is the only listed benchmark marked simultaneously as execution-based, repository-level, multilingual, refactoring-focused, averaging more than five modified files, and expert-curated.

Feature comparison of coding and refactoring benchmarks across execution, repository scope, multilinguality, refactoring focus, patch scale, and expert curation.
Feature comparison of coding and refactoring benchmarks across execution, repository scope, multilinguality, refactoring focus, patch scale, and expert curation.

The pipeline visualizes the selection funnel from repository search and commit extraction through Docker validation to expert filtering, problem rewriting, and verification. It explains why the released set is much smaller than the initial candidate pool.

Three-stage collection, environment-construction, and expert-curation workflow used to create the benchmark.
Three-stage collection, environment-construction, and expert-curation workflow used to create the benchmark.

3.3 Dataset composition

The 170 tasks average 11.4 source-patch files, 261.6 source LOC, and 8,179.5 source-patch tokens; test patches add 4.5 files and 185.5 LOC on average. The total average is 15.9 modified files, issue descriptions average 685.3 tokens, and the largest source patch modifies 182 files.

4 Experiments

Experiments compare six proprietary and open-weight models under mini-swe-agent and OpenHands. Evaluation uses all 170 instances, pre-built isolated Docker environments, a 300-step limit, a $10 per-instance cost limit, Pass@1 resolve rate, per-language resolve rates, agent steps, and average API cost.

4.1 Experimental setup

mini-swe-agent provides iterative file viewing, editing, search, and bash execution, whereas OpenHands provides sandboxed command execution and structured file-editing tools. Using the same scaffolds and limits across models is intended to support a controlled comparison of model and runtime-tooling effects.

4.2 Models

The proprietary models are Gemini-3-Pro, Claude Sonnet 4.6, and GPT-5.2; the open-weight models are GLM-5, Kimi-K2.5, and Qwen3.5. The selection spans model families, architectures, and training paradigms for a common large-scale refactoring workload.

5 Results and analysis

ProMax remains difficult under the reported setups: GPT-5.2 with OpenHands attains the highest overall resolve rate, 41.2%, below the 75%+ frontier performance cited for SWE-bench Verified. Results vary substantially with agent scaffold, programming language, and API cost, and no model leads in every language.

5.1 Main results

Under OpenHands, GPT-5.2 reaches 41.2%, Claude Sonnet 4.6 reaches 38.8%, and GLM-5 and Qwen3.5 each reach 36.5%; GPT-5.2 rises from 21.8% with mini-swe-agent. Language leaders are GPT-5.2 for Python (48.3%) and C (75.0%), Claude Sonnet 4.6 for TypeScript (53.6%) and Rust (63.6%), GLM-5 for Java (34.6%), Kimi-K2.5 for Go (43.5%), and Qwen3.5 for C++ (54.5%).

The table supplies the numerical basis for claims about scaffold sensitivity, language variation, and cost efficiency. It reports OpenHands GPT-5.2 at 41.2% overall and OpenHands GLM-5 at 36.5% for $0.24 per instance.

Overall and per-language resolve rates, average steps, and average costs for six models under mini-swe-agent and OpenHands.
Overall and per-language resolve rates, average steps, and average costs for six models under mini-swe-agent and OpenHands.

5.2 Agent behavior analysis

For Claude Sonnet 4.6 and Kimi-K2.5, agent file-modification distributions track gold patches on small tasks but under-cover larger patches: agents reach 90% cumulative coverage at about 10 modified files, whereas gold patches reach it at about 20. Failed runs also use more interaction rounds, consistent with incomplete cross-file propagation and unproductive edit–revert cycles.

The left cumulative distributions show that Claude Sonnet 4.6 and Kimi-K2.5 modify fewer files than the gold patches, especially for larger changes. The right panel separates passing and failing runs, showing that failures tend to require more interaction rounds.

Cumulative distributions comparing agent and gold-patch file counts, and interaction rounds for resolved versus unresolved trajectories.
Cumulative distributions comparing agent and gold-patch file counts, and interaction rounds for resolved versus unresolved trajectories.

5.3 Cost and efficiency analysis

Cost is not proportional to resolve rate: under OpenHands, Claude Sonnet 4.6 costs $4.77 per instance for 38.8%, GLM-5 costs $0.24 for 36.5%, and GPT-5.2 reaches 41.2% at $3.60. Qwen3.5 takes the most average steps under both scaffolds, 155.4 and 141.2, without leading overall; the paper also states that repositories use approved open-source licenses and discloses LLM assistance for curation, appendix classification, and paper editing.

6 Conclusion

The conclusion attributes the benchmark’s difficulty to its large cross-file scope and quality-controlled specifications and tests. It identifies incomplete propagation of changes across affected files as a central empirical limitation of current agents and reports that open-weight models can approach leading proprietary performance at much lower API cost.

The repository inventory documents the public code sources and licenses used in the benchmark, supporting the stated breadth of 70 repositories across seven languages.

Repositories and open-source licenses used in the benchmark, grouped by language.
Repositories and open-source licenses used in the benchmark, grouped by language.

Per-language statistics reveal uneven patch complexity: C++ and Java average 21.4 and 20.8 modified files, while C has the highest average LOC at 424.1. TypeScript has 28 instances but draws from only two repositories, a concentration caveat identified by the paper.

Per-language counts of repositories and instances, plus average files, LOC, and non-test files per task.
Per-language counts of repositories and instances, plus average files, LOC, and non-test files per task.

Every task has at least two category labels, and 46.5% have three or more, showing that the benchmark contains compound software-maintenance changes rather than isolated transformations.

Distribution of the number of task categories assigned to each instance.
Distribution of the number of task categories assigned to each instance.

The selected examples show the benchmark's range: the listed C++ instance changes 244 files, while the TypeScript example changes 27 files. Their summaries cover header unification, temporal resolution, configuration renaming, tensor APIs, stream buffering, provider formatting, and component APIs.

One representative refactoring instance per language, with repository, file count, LOC, and transformation summary.
One representative refactoring instance per language, with repository, file count, LOC, and transformation summary.

The co-occurrence matrix shows that API Interface Change and Refactoring Cleanup overlap in 79 instances. Bug Fix also co-occurs frequently with Refactoring Cleanup (48) and API Interface Change (33), showing that tasks commonly combine several maintenance concerns.

Co-occurrence matrix for the ten multi-label task categories.
Co-occurrence matrix for the ten multi-label task categories.

The skill labels identify cross-file reasoning in 169 instances and API semantics in 168, followed by interface-contract reasoning in 165 and pattern matching in 156. These annotations support the benchmark's emphasis on repository-wide coordination.

Multi-label distribution of reasoning skills required by benchmark instances.
Multi-label distribution of reasoning skills required by benchmark instances.

Appendix

  • The appendix lists the 70 source repositories and licenses, gives per-language patch statistics, and analyzes task categories and required skills. Refactoring Cleanup and API Interface Change occur in 66.5% and 65.3% of instances, while cross-file reasoning and API semantics are labeled for 99.4% and 98.8%; representative tasks range from a 27-file TypeScript API unification to a 244-file C++ header refactor.

Brief Thoughts

The benchmark’s contribution is the attempt to align task specifications, test suites, and behavior-preserving outcomes for large, multilingual refactoring tasks. Its 41.2% best result is conditional on the evaluated scaffolds and budgets, so it should be read as a benchmark outcome rather than a direct measure of all software-engineering capability; ongoing auditing of curation decisions, repository concentration, and test sufficiency remains important.