Open-world evaluations for measuring frontier AI capabilities
The case for long, messy, real-world tasks to evaluate AI agents
What is CRUX?
CRUX (Collaborative Research for Updating AI eXpectations) is a project for systematically conducting open-world evaluations: long-horizon tasks in real-world environments where success cannot be neatly specified or automatically graded. These evaluations complement benchmarks by testing what agents can do in settings that are too messy to standardize.
Each evaluation involves a long-horizon, real-world task; an agent scaffold that could in theory allow agents to solve the task; detailed log analysis; and a write-up that includes interpretations from collaborators with diverse perspectives. We plan to release new evaluations every 1–2 months. Our latest evaluation asks whether AI agents can conduct open-ended AI research.
The problem
Benchmarks saturate quickly and can’t capture the messiness of real-world tasks. Whatever is precise enough to benchmark is also precise enough to optimize for.
The approach
Open-world evaluations: small numbers of long-horizon tasks in real-world settings, with detailed log analysis and human intervention to elicit upper-bound capabilities.
CRUX
A collaborative project to systematically conduct open-world evaluations, with new experiments every 1–2 months across AI R&D, governance, and more.
As AI systems become more capable, evaluators must accept tradeoffs between evaluations that are constrained and scalable, and evaluations that are noisy and realistic. Open-world evaluations represent one end of this spectrum. Each approach has real strengths and real limitations.
Team
Core team
- Sayash KapoorPrinceton University
- Peter KirgisPrinceton University
- Andrew SchwartzPrinceton University, Cornflower Labs
- Stephan RabanserPrinceton University
- Arvind NarayananPrinceton University
Collaborators
- David AfricaUK AISI
- J.J. AllaireMeridian Labs
- Tilman BayerIndependent
- Rishi BommasaniStanford University
- Derrick Chan-SewIndependent
- Harry CoppockUK AISI
- Magda DuboisUK AISI
- Gillian HadfieldJohns Hopkins University
- Andy HallStanford University
- Sara HookerAdaption Labs
- Seth LazarAustralian National University, Johns Hopkins University
- Yue LingIndependent
- Nitya NadgirIndependent
- Steve NewmanGolden Gate Institute for AI
- Viet NguyenUniversity of Toronto
- Matilda OronaUC Berkeley
- Dimitris PapailiopoulosUW Madison, Microsoft Research
- Toby PilditchUK AISI
- Abhishek ShettyIndependent
- Shoshannah TekofskyAI Digest
- Helen TonerGeorgetown University (CSET)
- Cozmin UdudecUK AISI
- Konstantinos VoudourisUK AISI
Funding. We are grateful to Coefficient Giving, Schmidt Sciences, and the Princeton AI Lab for funding to support this project, and to OpenAI for providing API credits.
Cite
@misc{open_world_evals,
title = {Open-world evaluations for measuring frontier AI capabilities},
author = {Sayash Kapoor and Peter Kirgis and Andrew Schwartz and Stephan Rabanser and J.J. Allaire and Rishi Bommasani and Harry Coppock and Magda Dubois and Gillian Hadfield and Andy Hall and Sara Hooker and Seth Lazar and Steve Newman and Dimitris Papailiopoulos and Shoshannah Tekofsky and Helen Toner and Cozmin Ududec and Arvind Narayanan},
url = {https://arxiv.org/abs/2605.20520},
year = {2026}
}