Research
- Sep 2026 Environments must evolve with models Kennet
- Jul 2026 EnterpriseSWE Benchmark
- May 2026 COBOLBench: Measuring Frontier Agents on Enterprise COBOL Leaderboards
- May 2026 Introducing LegacySWE Essay
- Mar 2026 LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Paper
- Feb 2026 VideoBench: Evaluating Coding Agents on Source-Grounded Video Generation Leaderboards
- Feb 2026 FigmaBench: Throwing Frontier Models into Real-World Figma Leaderboards
- Dec 2025 Understanding Virality with Vision-Language Models Essay
- Oct 2025 Metaphi SimHub Essay
- Sep 2025 From Static Benchmarks to Dynamic Worlds Essay