Event date · · SWE Refactor Bench

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

FACT STATEMENT

SWE Refactor Bench is a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt. A three-stage evaluation protocol measures migration completeness and behavioural correctness: Migration Audit, Behavioural Tests, and Agentic Verification using 6 independent coding agents. Across 520 runs from 8 frontier models and 26 model-effort configurations, only 28 runs (5.4%) pass all three stages, 13 of the 20 tasks receive no accepted solution, and the best model is claude-opus-5.

What happened

SWE Refactor Bench introduces a benchmark of 20 whole-repository migrations covering four kinds of technical debt. It uses a three-stage evaluation: Migration Audit to verify the migration occurred, Behavioural Tests for correctness, and Agentic Verification with six independent coding agents to generate targeted tests for hidden behavioural differences. Across 520 runs from eight frontier models and 26 model-effort configurations, only 28 runs (5.4%) pass all stages; 13 of 20 tasks receive no accepted solution, and the best model is claude-opus-5.

Technical significance

The benchmark addresses the 'Blindness' hack where agents copy original implementations to pass tests without performing the migration. The three-stage protocol, especially Agentic Verification with six independent agents, aims to detect hidden behavioural differences. The low pass rate (5.4%) indicates that current frontier models struggle with long-horizon, whole-repository migrations, even with increased effort configurations.

Industry impact

This benchmark highlights a significant gap in coding agent capabilities for real-world software maintenance tasks like stack migrations. The failure of most models suggests that autonomous migration tools are not yet reliable for production use, which may slow enterprise adoption of AI for large-scale refactoring.

Decision value

For enterprises, the benchmark provides a way to evaluate coding agents for migration projects, potentially reducing risk in technology stack upgrades. However, current low success rates indicate that human oversight remains necessary, limiting immediate cost savings from autonomous migration.

What to watch

Observable next signals include whether model providers release improved versions targeting migration tasks, whether the benchmark is adopted by other research groups, and whether agent frameworks incorporate migration-specific strategies. The benchmark may drive research into long-horizon planning and verification for code agents.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.