LH-Bench
Long-horizon agents scored on subjective enterprise work by expert rubrics, curated artifacts and human preference.
Evaluate your agentPaper- Tasks
- 121
- Agents
- 7
- Results
- March 2026
Tasks
- Figma to code33Real Figma designs built into React and Tailwind, scored by a visual judge against the design and by expert preference.Results
- Programmatic content41Courses from source documents rendered as video, scored on citation grounding, educator rubrics and expert preference.Results
- Financial documents37Business bank statements read into labelled transactions, scored on extraction and categorisation against risk analysts' labels.No results yet
- Figma design10Design work inside a headless Figma on one design system, checked for completeness and consistency.No results yet
Leaderboard
Figma to codeMetric Output scoreTasks 33
- 1GPT-5.2 ProCodex4.27
- 2Claude Opus 4.6Claude Code4.19±0.280
- 3GPT-5.2Codex3.94±0.360
- 4Claude Opus 4.5Claude Code3.88
- 5Gemini 3.1 ProGemini CLI3.73±0.440
- 6Claude Sonnet 4.5Claude Code3.66
- 7Gemini 3 ProGemini CLI3.59
7 agents
A visual judge scores the built page against the design, 1 to 5, on eight weighted rubrics. 95% bootstrap intervals over tasks where the run count allows.
Programmatic contentMetric Artifact scoreTasks 41
- 1Claude Opus 4.6Claude Code0.612±0.069
- 2Gemini 3.1 ProGemini CLI0.526±0.063
- 3GPT-5.2Codex0.478±0.044
3 agents
A visual judge scores each rendered chapter against rubrics written by educators, normalised to 0 to 1. Course-level means over 41 courses and 183 chapters, 95% bootstrap intervals.
The idea
Real enterprise work is subjective, and a unit test cannot score it. LH-Bench scores long-horizon agents with expert-grounded rubrics for the process, curated artifacts for the output and pairwise human preference to confirm the ranking.
Environments
- FigmaA headless Figma runtime with the Plugin API and MCP tools.
- Financial document reasoningBank statements and filings as PDFs, every transaction labelled by risk analysts.
- Content generationProgrammatic video from source documents, judged by expert educators.