EnterpriseSWE
Long-horizon software engineering on real enterprise systems, verified by execution.
Evaluate your agent- Tasks
- 27
- Agents
- 2
- Results
- August 2026
Tasks
- Mainframe27Tickets on a real telecom billing system, fixed across the programs and jobs they touch and verified by running them.Results
Leaderboard
MainframeMetric pass@1Tasks 20
- 1Claude Fable 5Claude Code25.0%11.2% to 46.9%
- 1kimi-k3Terminus 225.0%11.2% to 46.9%
2 agents
One attempt per task, verdict by the verifier on the real system, 95% Wilson interval. Both agents are scored on the same 20 of the 27 tasks; seven are left out because the Claude Fable 5 first attempts on them ended in harness errors.
The idea
Frontier agents write code well and still fail inside the systems that run banks, insurers and telecoms, where decades of COBOL, JCL and DB2 spread one business rule across hundreds of programs. EnterpriseSWE evaluates agents on those systems, mainframe first, with a verifier that runs the changed programs and checks the business outcome.