Global Tech Giants Are Betting on AI for Science—At Last, There’s a Unified Standard
In August this year, Google Chief Scientist Jeff Dean announced his departure to found Discovery Loop, an AI for Science (AI4S) company.
That same year, tech giants including OpenAI, Anthropic, and NVIDIA also announced their entry into the AI4S space.
These “cross-industry” players share a highly consistent goal: not merely to build tools that assist with research, but to create “AI scientists” capable of autonomously completing the full research loop of “hypothesis → experiment design → execution and validation → iterative optimization.”
And it is not just the tech giants. Many companies entered the field even earlier. Founded in 2024, Chinese company Deep Principle is one of the earliest companies in the world to bet on autonomous AI research.
Behind this sector lies a vast “gold mine.” Once AI can achieve autonomous, closed-loop research, it can evolve from “selling software” into the productivity foundation for trillion-dollar physical industries such as new materials and pharmaceuticals.

At the national-strategy level, this sector is also receiving considerable attention. AI-driven scientific discovery has been incorporated into China’s latest strategy for becoming a global science and technology powerhouse.
The hotter the industry becomes, the harder it is to avoid one fundamental question: How do we determine whether these “AI scientists” are actually performing well?
The Old Yardsticks Are No Longer Enough
For years, the standard yardstick for measuring AI’s research capabilities has been closed-book knowledge exams such as HLE.
These exams are essentially tests that evaluate only the final answer. They focus on the result, not on how the answer was produced.
AI research tools that score highly on these exams often reveal their limitations as soon as they are asked to conduct a real experiment: they may be able to cite literature fluently, yet fail to stop and examine or reflect on anomalous data. They may also propose experiments that sound perfectly reasonable but cannot actually be carried out.
On August 19, a paper titled Measuring AI Scientists: From Exams to Discovery was published. It introduced, for the first time, a systematic and comprehensive evaluation framework for “AI’s ability to conduct real scientific research,” officially addressing a long-standing gap in the field’s evaluation standards.
The paper was jointly produced by Chinese AI4S company Deep Principle, Microsoft Research, Stanford University, FutureHouse, Edison Scientific, the Allen Institute for AI, the University of California, Berkeley, and other leading research and industry institutions worldwide.

The framework’s academic authority is underscored by the participation of leading scholars in AI4S, including Kristin Persson, founder of the world’s largest open materials-data initiatives, Materials Genome and Materials Project; Andrew White, founder of FutureHouse and Edison Scientific and a key author of the recent Co-Scientist paper in Nature; Peter Jansen, a core developer of ScienceWorld, a benchmark for scientific agents; Yuanqi Du, first author of a 2023 Nature review of AI4S; and Ludwig Schmidt, the final corresponding author of Terminal Bench.
As the only Chinese industry representative, Deep Principle contributed first-hand experience accumulated through real-world R&D pipelines, giving the framework practical applicability.

The paper even carries the signature of Nobel Prize-winning chemist Frances Arnold. Her participation—despite not being known primarily for AI and rarely appearing publicly in the AI4S field—signals the scientific community’s growing recognition of the importance of this direction.
In 2026, as players of all kinds—from tech giants to startups—rushed into AI research, competition across the sector intensified. The emergence of this evaluation framework at this particular moment was no coincidence.
How to Systematically and Comprehensively Evaluate AI’s Research Capabilities
The paper’s central innovation is its proposal of the discovery episode evaluation framework.
At its core, the framework takes the form of a comprehensive assessment that records every decision AI makes throughout the entire research cycle, ultimately providing an overall score for the scientific rigor, validity, and effectiveness of the complete trajectory.
This approach is fundamentally different from the exam-oriented logic used previously.
Organized around the three stages of the research loop—scientific hypothesis, execution of the proposed approach, and interpretation of experimental results—the framework establishes corresponding evaluation dimensions for each stage. It also concludes with a “full-process closed-loop test” to measure AI’s overall ability to produce genuine scientific discoveries.

△The paper highlights the difference between knowledge evaluation and discovery evaluation, and defines the key elements included in a “discovery episode”
The framework also redefines the value of failed data: failed experiments are precisely the core assets for training and evaluating AI’s research capabilities.
Human scientists can learn from failure and avoid taking unnecessary detours. AI conducting research should likewise learn to extract insights from failure.
In addition, the assessment sets high standards for the authenticity and independence of AI-generated research results, ensuring that experimental outcomes are “genuine” new discoveries rather than “self-proclaimed achievements” that cannot be empirically verified.
Deep Principle’s MIRA Is a Real-World Example of This Framework
Even before the publication of this standard-setting paper, leading players had already implemented the closed-loop logic in their products. Deep Principle’s MIRA platform is one of the most representative industry examples.
The role of ordinary research AI is that of an “answering tool”: researchers pose a clearly defined question, and the AI returns a corresponding answer.
MIRA’s design philosophy is entirely different from that of conventional research AI. It builds a complete closed-loop system spanning “hypothesis generation → computational simulation → experimental validation → knowledge retention,” and replicates the complete collaboration model of a real research team through a three-layer architecture.
- The upper layer consists of a multi-agent team with specialized roles and coordinated collaboration, capable of centrally planning and scheduling the research process;
- The middle layer consists of dual execution engines connecting high-performance computing with automated laboratories, enabling coordination between computational and experimental workflows;
- The bottom layer is a persistent, reusable research memory system that preserves each project’s data and conclusions from trial and error, turning them into a foundation for subsequent research.

△The MIRA AI Scientist interface in action
The industrial value of this design has already been validated through public benchmarks.
MIRA ranked first in both the Research Claw Benchmark, which evaluates chemistry, energy, and materials research, and the Science Agent Arena benchmark for drug discovery. Its average cost per completed task is also among the lowest in the industry, giving it an advantage on both performance and cost.

△MIRA’s results in the Research Claw Benchmark and other evaluations

△Science Agent Arena drug discovery leaderboard: MIRA ranked first with an overall score of 81.1%
Deep Principle’s previously released React-OT model can calculate chemical reaction transition states in 0.4 seconds, establishing the platform’s foundation of high-precision computational capabilities.
Building on specialized models like this, MIRA further connects the entire research workflow. This evolution mirrors the broader AI4S industry’s transition from tools to platforms and from isolated capabilities to closed-loop systems.
As a core industry co-creator of this evaluation framework, Deep Principle has leveraged MIRA and its proprietary high-throughput experimental platform to establish closed-loop computational and experimental workflows across multiple self-developed research pipelines in lithium batteries, industrial coolants, and new energy materials. AI leads the entire research cycle, generating real-world closed-loop data and enabling continuous mutual validation and iteration between the platform and the evaluation framework.
Comprehensive research data spanning computational and experimental laboratories is a critical asset for helping AI quickly understand the boundaries of hypotheses and avoid repeated trial and error. It is also expected to become key infrastructure for the next stage of AI for Science.
Conclusion
The release of this standard framework marks a milestone for the AI for Science sector, providing a clear yardstick for all the players currently racing forward.
But generating the data needed to validate this new yardstick is no easy task. Only AI-native companies such as Deep Principle and Edison Scientific, with end-to-end R&D capabilities spanning chemistry, materials science, and biology, are likely to achieve it.
In the first half of the race, the industry competed over “who knows more.” Tech giants competed on foundation models, parameter counts, and exam scores, using increasingly difficult question banks to demonstrate the depth of their models’ knowledge.
These efforts laid the foundation for AI research, but they have also gradually reached a ceiling.
In the second half, the competition will be about “who can actually make something.”
Since the beginning of this year, from overseas tech giants entering the field one after another, to Jeff Dean founding Discovery Loop, to the closed-loop research practices of domestic companies such as Deep Principle, the entire industry has been moving toward one goal: enabling AI to truly complete the full loop from hypothesis to discovery.
Today, the evaluation criteria for AI research are increasingly aligned with the logic of science itself. The day when AI becomes a genuinely trustworthy source of research productivity is drawing ever closer.
For the complete design details, testing scenarios, and quantitative standards of the “discovery episode” evaluation framework, see the original paper:
Measuring AI Scientists: From Exams to Discovery