Metaphi AI

LH-Bench

Long-horizon agents scored on subjective enterprise work by expert rubrics, curated artifacts and human preference.

Evaluate your agentPaper
Tasks
121
Agents
7
Results
March 2026

Tasks

  • Figma to code33Real Figma designs built into React and Tailwind, scored by a visual judge against the design and by expert preference.Results
  • Programmatic content41Courses from source documents rendered as video, scored on citation grounding, educator rubrics and expert preference.Results
  • Financial documents37Business bank statements read into labelled transactions, scored on extraction and categorisation against risk analysts' labels.No results yet
  • Figma design10Design work inside a headless Figma on one design system, checked for completeness and consistency.No results yet

Leaderboard

Figma to codeMetric Output scoreTasks 33
  1. 1GPT-5.2 ProCodex4.27
  2. 2Claude Opus 4.6Claude Code4.19±0.280
  3. 3GPT-5.2Codex3.94±0.360
  4. 4Claude Opus 4.5Claude Code3.88
  5. 5Gemini 3.1 ProGemini CLI3.73±0.440
  6. 6Claude Sonnet 4.5Claude Code3.66
  7. 7Gemini 3 ProGemini CLI3.59
7 agents

A visual judge scores the built page against the design, 1 to 5, on eight weighted rubrics. 95% bootstrap intervals over tasks where the run count allows.

Programmatic contentMetric Artifact scoreTasks 41
  1. 1Claude Opus 4.6Claude Code0.612±0.069
  2. 2Gemini 3.1 ProGemini CLI0.526±0.063
  3. 3GPT-5.2Codex0.478±0.044
3 agents

A visual judge scores each rendered chapter against rubrics written by educators, normalised to 0 to 1. Course-level means over 41 courses and 183 chapters, 95% bootstrap intervals.

The idea

Real enterprise work is subjective, and a unit test cannot score it. LH-Bench scores long-horizon agents with expert-grounded rubrics for the process, curated artifacts for the output and pairwise human preference to confirm the ranking.

Environments

  • FigmaA headless Figma runtime with the Plugin API and MCP tools.
  • Financial document reasoningBank statements and filings as PDFs, every transaction labelled by risk analysts.
  • Content generationProgrammatic video from source documents, judged by expert educators.