Research
We believe there is significant under exploration in data and environment design research. We publish our work on our experiments on novel data recipes and environment scaling.
- EssaySep 2026Scaling compute on tacit knowledgeWe propose a multi-loop harness that consults human experts as it designs environments and builds tasks, allowing us to scale task construction beyond the hours experts can spend hand authoring.Read essay
- PaperMar 2026LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise TasksLong-horizon agents scored on subjective enterprise work by expert rubrics, curated artifacts and human preference.Read paper
- EssayDec 2025Understanding Virality with Vision-Language ModelsA rubric-based vision-language model framework for short-form edutainment evaluation.Read essay
- EssayOct 2025Metaphi SimHubAn interactive simulation environment for multi-step code generation.Read essay
- EssaySep 2025From Static Benchmarks to Dynamic WorldsThere is no "user" in the evaluation loop.Read essay