Skip to content
Main Site News Console

Did Google’s AI agent really build an operating system for just $916?

· Normal Tech (AI Snake Oil) Translated
分析实践

By Stephan Rabanser, Sayash Kapoor, Rishi Bommasani, Andrew Schwartz, and Arvind Narayanan

Earlier this week, at Google’s developer conference, the company unveiled its latest model, Gemini 3.5 Flash, along with a new agent application, Antigravity 2.0. To demonstrate the capabilities of this new agent system, Google claimed that a team of agents had built a complete operating system. Reportedly, the project required just a single prompt, cost approximately $900 in API fees, and was completed by dozens of sub-agents working together.

Does this mean AI can now build complex software at low cost? Not so fast:

  • The claim that it was built with a “single prompt” is misleading. The blog post says that the operating system was built using a single prompt. But midway through the article, Google reveals that the prompt “eventually grew to thousands of lines.” How many attempts did it take to generate that prompt? How specific were the instructions given to the agents? Without these crucial details, it is difficult to determine whether the real secret was a better model or simply investing more effort in the prompt. In addition, the run was conducted using a scaffold1 that included specialized roles, task delegation to sub-agents, and agents for detecting and preventing cheating. In its announcement, Google presents this scaffold as a product feature. But we do not know whether the scaffold was overfit to the task of “building an operating system from scratch,” or whether it would perform equally well on other complex software engineering tasks.

  • Google’s article does not clearly explain what counted as human intervention. The article says that the final run to develop the operating system “did not require additional guidance or correction from humans.” But it does not define this standard. The article describes infrastructure for terminating and restarting agents that became stuck. It also mentions that, during an earlier run, the agents appeared to cheat, after which the team added anti-cheating measures and reran the task. But the article does not include these trial runs in its methodological account. Nor does it clearly state whether any agents escalated issues to humans, whether the final run required humans to restart, approve, or repair anything, or how many attempts were made before the agents succeeded.

  • The article provides no analysis to determine whether the agents wrote the code from scratch or copied existing code from the internet. To Google’s credit, the blog post notes that toy operating systems are common undergraduate projects and that publicly available implementations are easy to find. The article itself also raises the concern that the agents may simply have reproduced existing information rather than building an operating system from scratch. But it does not address this concern—it provides neither similarity analysis nor log analysis to check whether the agents copied existing code. Even in the absence of direct copying, building an operating system may be relatively easy for agents because they have already memorized relevant patterns from their training data. Thus, this does not provide sufficient evidence of an agent’s ability to create novel software.

  • Google did not release the lengthy prompt, the code written by the agents, or the run logs, making it impossible to independently evaluate these claims. If the source code or agent logs were released, independent researchers could assess the quality of the resulting system and answer questions such as whether the agents copied existing code. The blog post contains only a short video documenting one segment of the development process, along with an overview of the experiment as a whole.

On the other hand, the blog post does report the exact cost of building the operating system ($916.92), as well as the total token budget (2.6 billion tokens in all). This information provides useful context, and that is worth recognizing. Many of the evaluations we have previously examined disclosed no costs at all, making it difficult to compare their central findings with those of other evaluations.

Still, Google’s blog post is essentially a press release. We understand that it is unrealistic to expect it to meet the standards of scientific rigor. Evaluations like this one—in which a long-horizon task is carried out in the real world, run only once, and described by the experimenters after the fact—have become increasingly common. Because many of these evaluations are conducted by AI companies, it is easy to dismiss the entire category as self-aggrandizing.

But that would be a mistake. We call this emerging paradigm open-world evaluations, and we discuss the trend in a recent paper and accompanying blog post. The key point is that we believe open-world evaluations require an entirely new set of methodological norms. When implemented properly, they can provide a valuable perspective that benchmark evaluations cannot.

Google’s experiment adds to the evidence that agents—or teams of agents—can handle certain types of tasks autonomously or nearly autonomously over extended periods, continuing to make progress without getting stuck or confused. As we argue in our paper, benchmark evaluations are in fact difficult to carry out for these kinds of tasks, for many reasons including cost. This is therefore an opportune time for independent evaluators in academia, nonprofit organizations, and government to step in and bring rigor and credibility to open-world evaluations—qualities that are unlikely to feature in AI vendors’ own claims.

Footnotes

  1. A scaffold is a layer of code, prompts, and tools built around an AI model to enable it to act autonomously, including functions for managing memory, accessing tools, and interacting with its environment. For example, Claude Code is the scaffold that enables Anthropic’s Claude model to operate as a coding agent. ↩