VideoBench: Evaluating Coding Agents on Source-Grounded Video Generation
A benchmark measuring AI agents on video, animation, and presentation generation from curated data rooms. 183 chapters, 41 courses.
We are on a mission to serve mission-critical industries with frontier intelligence.
Our research focus is to build self-improving agentic systems.
The continual learning harness for agents in production.
Environments where agents learn mission-critical enterprise engineering.
We evaluate frontier coding systems on a 100-task public release drawn from real enterprise COBOL maintenance environments.
A benchmark measuring AI agents on video, animation, and presentation generation from curated data rooms. 183 chapters, 41 courses.
A benchmark measuring AI agents on production Figma-to-code conversion through API interaction, design hierarchy extraction, and iterative deployment.
How principles from autonomous vehicle simulation can transform code generation agent training and evaluation.
We partner with frontier labs and frontier enterprises on environments and OTS datasets.