Skip to content
Main Site News Console

AI Agents Still Can’t Conduct Open-Ended AI Research

· Normal Tech (AI Snake Oil) Translated
分析实践

The goal of leading AI labs is recursive self-improvement (RSI): automating AI research using AI agents. RSI also underpins predictions that AI progress will explode. How can we assess how close we are to this milestone?

One approach is to use benchmarks to test whether agents can conduct AI research. Given the AI community’s emphasis on benchmarks, they have been the primary way to evaluate progress toward RSI. Over the past year, many such evaluations have found that agents can now make progress on tasks where success is easy to verify, prompting speculation that RSI may be just around the corner.

But while these evaluations are useful, they are limited to narrow, verifiable tasks. AI research may be much more open-ended. Success is often not immediately clear or verifiable; to make progress, researchers need to test promising hypotheses, backtrack, or consider new and unconventional approaches. How should we evaluate agents’ ability to conduct open-ended AI research? We took a first step toward answering this question in a new paper.

We worked with the authors of two unpublished AI papers and asked them to formulate the papers’ primary research questions. We then tasked frontier AI agents with conducting research to answer those questions, giving them several thousand dollars’ worth of API credits and compute, along with six days of real-world runtime. The original paper authors reviewed the papers written by the agents.

The authors explicitly rejected both agent-written papers. To better understand these results, our team spent more than 100 hours analyzing the agents’ logs. Our main conclusions are as follows:

  1. Agents lack the judgment required for open-ended research. Although agents proposed several research directions that impressed expert reviewers, they quickly rejected those directions based on low-quality or synthetic data.
  2. Agents do not understand the resources available to them. At the end of both experiments, less than 50% of the API budget had been used, and several hours remained before the deadline—even though the agents could monitor resource usage and were encouraged to use the full budget.
  3. Agents cannot respond creatively to feedback. Although the agents’ own AI reviews identified many of the issues later pointed out by expert reviewers, the agents did not solve those issues creatively. When faced with negative feedback, they merely added qualifications to existing findings and continued pursuing unpromising research directions.
  4. Agents cannot backtrack effectively. They abandoned their most ambitious research goals within the first day of the experiment, and neither agent fundamentally changed its approach afterward.
  5. Agents did not follow specific instructions. They ignored explicit requirements concerning how much time to spend in the exploration phase, how frequently to seek feedback from AI review tools, and the maximum paper length.

Two years ago, we published a benchmark examining whether agents could be used to improve research reproducibility. Since then, we have wanted to assess AI’s ability to conduct open-ended research. But we wanted to design the methodology more rigorously. The approach was originally proposed by several UK Artificial Intelligence Safety Institute (UK AISI) coauthors of the paper and further developed by the core team at Princeton University.

We call this approach “shadow evaluations” because the agents work alongside the original research. In addition to the two of us, the core team included Peter Kirgis, Andrew Schwartz, and Stephan Rabanser. The full list of authors and affiliations appears at the end of this article.

Shadow evaluations offer important advantages: they allow us to test agents on research findings that the agents have not been trained on and cannot access online. They also allow experts who spent months answering the relevant questions to evaluate the agents’ outputs.[^1]

But shadow evaluations also have inherent limitations. Expert reviewers know that the papers were generated by AI, and may therefore favor the approaches they themselves took over those taken by the agents. Because we conduct an in-depth evaluation of each paper, the sample size is small (in this study, we used only two papers). In addition, this type of evaluation inevitably requires researchers to exercise considerable flexibility in its design, execution, and interpretation.

Indeed, we are known for holding a particular position in discussions of recursive self-improvement and superintelligence. This may influence how we conduct our research. The paper includes a section discussing our potential biases and how we addressed them. We sought out collaborators who did not fully share our prior views and explicitly present the resulting disagreements.[^2] In future evaluations, we would be interested in making “adversarial collaborators” part of the core team.

Implications for an Explosion in AI Progress

Our results suggest that open-ended research remains challenging for frontier AI agents. However, these findings are still preliminary, and we are working to address their limitations—for example, by expanding the sample size, testing new models, and potentially improving the scaffold. But if these findings hold up under further scrutiny, what are their implications?

First, we need to understand the extent to which progress in frontier AI—and RSI—can be achieved simply through continued hill climbing on verifiable tasks. Our view is that although faster progress on narrow tasks, such as improving efficiency, is certainly possible, we do not expect it to produce broad RSI or explosive progress. Nevertheless, we plan to closely track how agents’ capabilities on verifiable tasks affect the trajectory of AI progress.

Second, we need to measure how quickly the limitations currently facing agents conducting open-ended research—such as their lack of creativity and judgment—can be overcome, for example through more targeted training and improvements to the scaffold. We plan to continue conducting shadow evaluations regularly to help answer this question.

Finally, even if these limitations can be overcome, other bottlenecks may exist that could slow the pace of AI progress.