Open-world evaluations for measuring frontier AI capabilities

The case for long, messy, real-world tasks to evaluate AI agents

What is CRUX?

CRUX (Collaborative Research for Updating AI eXpectations) is a project for systematically conducting open-world evaluations: long-horizon tasks in real-world environments where success cannot be neatly specified or automatically graded. These evaluations complement benchmarks by testing what agents can do in settings that are too messy to standardize.

Each evaluation involves a long-horizon, real-world task; an agent scaffold that could in theory allow agents to solve the task; detailed log analysis; and a write-up that includes interpretations from collaborators with diverse perspectives. We plan to release new evaluations every 1–2 months. Our latest evaluation asks whether AI agents can conduct open-ended AI research.

The problem

Benchmarks saturate quickly and can’t capture the messiness of real-world tasks. Whatever is precise enough to benchmark is also precise enough to optimize for.

The approach

Open-world evaluations: small numbers of long-horizon tasks in real-world settings, with detailed log analysis and human intervention to elicit upper-bound capabilities.

CRUX

A collaborative project to systematically conduct open-world evaluations, with new experiments every 1–2 months across AI R&D, governance, and more.

As AI systems become more capable, evaluators must accept tradeoffs between evaluations that are constrained and scalable, and evaluations that are noisy and realistic. Open-world evaluations represent one end of this spectrum. Each approach has real strengths and real limitations.

Simple
Complex
Single-turn Q&A
+ Broad knowledge assessment, scalable, reproducible
Multiple-choice format is artificial; users rarely interact with models this way. Increasingly saturated for frontier models.
Open-ended chat
+ Captures nuance in free-form responses
Limited to single-turn or short interactions. Cannot measure long-horizon planning or tool use.
Outcome-only agent benchmarks
+ Tests agent performance on real, well-defined tasks
Only measures whether the task was completed, not how. Most passing SWE-Bench solutions are not accepted by maintainers.
Agent benchmarks with log analysis
+ Examines how agents succeed or fail, uncovering reward hacking
Still operates in sandboxed environments with predefined tasks. Cannot capture real-world messiness.
Open-world evaluations
+ Long-horizon, real-world tasks that elicit upper-bound capabilities
Not reproducible or standardized. Hard to compare across agents. Success criteria can be blurry.

Team

Core team

  • Sayash Kapoor
    Princeton University
  • Peter Kirgis
    Princeton University
  • Andrew Schwartz
    Princeton University, Cornflower Labs
  • Stephan Rabanser
    Princeton University
  • Arvind Narayanan
    Princeton University

Collaborators

  • David Africa
    UK AISI
  • J.J. Allaire
    Meridian Labs
  • Tilman Bayer
    Independent
  • Rishi Bommasani
    Stanford University
  • Derrick Chan-Sew
    Independent
  • Harry Coppock
    UK AISI
  • Magda Dubois
    UK AISI
  • Gillian Hadfield
    Johns Hopkins University
  • Andy Hall
    Stanford University
  • Sara Hooker
    Adaption Labs
  • Seth Lazar
    Australian National University, Johns Hopkins University
  • Yue Ling
    Independent
  • Nitya Nadgir
    Independent
  • Steve Newman
    Golden Gate Institute for AI
  • Viet Nguyen
    University of Toronto
  • Matilda Orona
    UC Berkeley
  • Dimitris Papailiopoulos
    UW Madison, Microsoft Research
  • Toby Pilditch
    UK AISI
  • Abhishek Shetty
    Independent
  • Shoshannah Tekofsky
    AI Digest
  • Helen Toner
    Georgetown University (CSET)
  • Cozmin Ududec
    UK AISI
  • Konstantinos Voudouris
    UK AISI

Funding. We are grateful to Coefficient Giving, Schmidt Sciences, and the Princeton AI Lab for funding to support this project, and to OpenAI for providing API credits.

Cite

@misc{open_world_evals,
  title = {Open-world evaluations for measuring frontier AI capabilities},
  author = {Sayash Kapoor and Peter Kirgis and Andrew Schwartz and Stephan Rabanser and J.J. Allaire and Rishi Bommasani and Harry Coppock and Magda Dubois and Gillian Hadfield and Andy Hall and Sara Hooker and Seth Lazar and Steve Newman and Dimitris Papailiopoulos and Shoshannah Tekofsky and Helen Toner and Cozmin Ududec and Arvind Narayanan},
  url = {https://arxiv.org/abs/2605.20520},
  year = {2026}
}