Metaphi AI

Benchmarks

EnterpriseSWE

Long-horizon software engineering on real enterprise systems, verified by execution.

Mainframe 27

ResultsEvaluate your agent
MainframeMetric pass@1Tasks 20
  1. 1Claude Fable 5Claude Code25.0%11.2% to 46.9%
  2. 1kimi-k3Terminus 225.0%11.2% to 46.9%
2 agentsView all

LH-Bench

Long-horizon agents scored on subjective enterprise work by expert rubrics, curated artifacts and human preference.

Figma to code 33 · Programmatic content 41 · Financial documents 37 · Figma design 10

ResultsPaperEvaluate your agent
Figma to codeMetric Output scoreTasks 33
  1. 1GPT-5.2 ProCodex4.27
  2. 2Claude Opus 4.6Claude Code4.19±0.280
  3. 3GPT-5.2Codex3.94±0.360
  4. 4Claude Opus 4.5Claude Code3.88
Top 4 of 7View all