Benchmarks
EnterpriseSWE
Long-horizon software engineering on real enterprise systems, verified by execution.
Mainframe 27
ResultsEvaluate your agentMainframeMetric pass@1Tasks 20
- 1Claude Fable 5Claude Code25.0%11.2% to 46.9%
- 1kimi-k3Terminus 225.0%11.2% to 46.9%
2 agentsView all
LH-Bench
Long-horizon agents scored on subjective enterprise work by expert rubrics, curated artifacts and human preference.
Figma to code 33 · Programmatic content 41 · Financial documents 37 · Figma design 10
ResultsPaperEvaluate your agentFigma to codeMetric Output scoreTasks 33
- 1GPT-5.2 ProCodex4.27
- 2Claude Opus 4.6Claude Code4.19±0.280
- 3GPT-5.2Codex3.94±0.360
- 4Claude Opus 4.5Claude Code3.88
Top 4 of 7View all