Skip to content
Main Site News Console

Open-World Benchmark for Measuring Frontier AI Capabilities

· Normal Tech (AI Snake Oil) Translated
分析实践

This article is 8,000 words long—our latest collaborative paper on an emerging type of AI evaluation. The paper has also been published in PDF format here.

Abstract: AI models have begun to saturate most major benchmarks. But does this mean they can build and deliver a real product, conduct a scientific experiment end to end, or navigate a government bureaucracy? Researchers have begun testing AI in these real-world environments. We call this type of evaluation “open-world evaluations.” This article defines open-world evaluations, summarizes the lessons learned so far, and proposes best practices for conducting them.

We will also introduce CRUX, a collaborative project involving 17 researchers from academia, government, civil society, and industry. Its goal is to regularly evaluate frontier AI capabilities through open-world evaluations. In our first experiment, an AI agent built and published an iOS app to the App Store, making only two mistakes, one of which required human intervention. This result offers early evidence that AI may already possess potentially useful capabilities. More importantly, it also provides an early warning about the risks of AI-driven app-store spam—we disclosed this result to Apple one month before publishing the paper.

We hope to conduct similar experiments in other real-world domains to identify risks early. This will be one of our main empirical research programs over the coming year.

Authors: Sayash Kapoor、Peter Kirgis、Andrew Schwartz、Stephan Rabanser、J.J. Allaire、Rishi Bommasani、Magda Dubois、Gillian Hadfield、Andy Hall、Sara Hooker、Seth Lazar、Steve Newman、Dimitris Papailiopoulos、Shoshannah Tekofsky、Helen Toner、Cozmin Ududec、Arvind Narayanan


How should we track and forecast the capabilities of AI? Today, the mainstream answer in the AI community is benchmarking. For example, METR’s time horizon chart has been used by policy analysts, industry leaders, and organizations researching AI risk to argue that AI capabilities are improving rapidly.

But benchmarks can both overestimate and underestimate progress. To turn a task into a benchmark, it must be specified precisely and its outcome must be verifiable automatically. The problem is that anything precise enough to be included in a benchmark is also precise enough to be specifically optimized for, allowing AI agents to perform well on such tasks. Conversely, low accuracy on a benchmark may result from incidental failures—for example, encountering a CAPTCHA on a website—even when the agent itself is capable of solving the underlying task.

To address these limitations, many researchers have begun turning to a new type of evaluation: longer-running, messier real-world evaluations that go beyond benchmarks. Nicholas Carlini of Anthropic used a Claude agent to build a C compiler capable of compiling the Linux kernel. Anthropic and Andon Labs also designed a free-form experiment in which Claude was tasked with running a small shop in a company office. Benchmarks typically contain dozens of tasks and are evaluated automatically, whereas open-world evaluations involve smaller sample sizes, often require human intervention, and are assessed in an open-ended manner—for example, by analyzing agent logs.

It is easy to regard these evaluations as insufficiently scientific: each evaluation has a sample size of one, and they lack standardization and reproducibility. Despite these limitations, we believe such evaluations are important for gathering evidence about AI capabilities. They can provide early warnings about emerging capabilities, inform efforts to build societal resilience, help evaluators identify blind spots in existing benchmarks, and give businesses a clearer understanding of which tasks AI systems may soon be able to perform, thereby informing strategic decisions about AI. We call them open-world evaluations.

This article conceptualizes open-world evaluations, reviews past examples to identify best practices and potential challenges in conducting them, and introduces the CRUX project, which aims to conduct new open-world evaluations on a regular basis. Here are our main points:

  • Open-world evaluations are an important emerging category of AI evaluation. As AI systems become more capable, evaluations designed to elicit frontier capabilities must also become more sophisticated. Open-world evaluations are the latest in a range of increasingly complex evaluation types. We surveyed 10 representative open-world evaluations conducted over the past year to identify best practices and key lessons.

  • CRUX (Collaborative Research for Updating AI eXpectations) is an attempt to conduct open-world evaluations systematically. Team members come from government, academia, and nonprofit organizations. Many of them have led open-world evaluations before, and they hold a range of views about how AI will develop in the future. Our goal is to provide empirical evidence about the current capabilities of AI systems, even when those capabilities are currently expensive to use, while also providing early warnings about capabilities that may soon become widespread. We plan to publish new open-world evaluations regularly.

  • In CRUX’s first experiment, we asked an AI agent to develop a simple iOS app and publish it to the App Store. Many benchmarks evaluate an agent’s ability to write code. But publishing an iOS app involves many other steps: signing the app, publishing a privacy policy on the web, completing Apple’s forms, and getting the app through the review process. Rather than focusing on the agent’s ability to write code, we were more interested in whether it could satisfy the real-world requirements involved in publishing an app. We therefore asked it to build a simple app and complete the iOS App Store submission process.

  • The agent successfully completed the task after making two mistakes, one of which required human intervention (forgetting where the correct credentials were stored and fabricating a nonexistent phone number during the App Store review process). The total cost of developing and publishing the app was approximately $1,000. The app is now live on the iOS App Store. We believe the cost could have been much lower: app development and submission cost only $25; the vast majority of tokens were spent monitoring the app’s status. We contacted Apple one month before publishing this article to disclose the experimental results. App-store operators should prepare for and strengthen their governance of spam-app submissions, as they may soon see thousands of apps submitted autonomously by agents.

  • How can we improve open-world evaluations? What comes next? To make open-world evaluations more useful, evaluators should clearly specify what kinds and how much human intervention are permitted; publish the logs collected while the agent solves the task; and analyze those logs to report what actions the agent took while solving the task. In future CRUX evaluations, we will assess AI R&D automation, AI governance, and many other domains.

Open-world evaluations are an important emerging category of AI evaluation

This section defines open-world evaluations and surveys the emerging ecosystem of this evaluation category to distill its successes and limitations. We will discuss the areas in which open-world evaluations can overcome some of the blind spots of benchmarks. We believe that as AI systems become more capable, the evaluations designed to elicit frontier