Skip to content
Main Site News Console

AI4S Enters the “Project Era”: Zidong Taichu Takes AI from Tasks to Projects

· 量子位
国内AI

AI for Science Is Shifting from Tool-Level to System-Level Capabilities

Wen Le, reporting from Aofeisi
QbitAI | WeChat Official Account: QbitAI

AI for Science is entering a new watershed.

Over the past few years, with the rapid development of foundation models and agent technologies, AI has gradually expanded from literature retrieval, data analysis, and code generation into research processes such as simulation and experimental analysis.

“Can AI help scientists conduct research?” is no longer the only question.

A new question is emerging:

Can AI truly take on an entire research project, rather than merely completing individual research Tasks?

This reflects a shift in the capability paradigm of AI4S.

In the past, AI primarily served as a tool or task-oriented agent. Once researchers defined the objectives, boundaries, and methodology, AI would complete a series of relatively independent Tasks.

But real-world research is not simply a collection of tasks.

A Project often begins with a vague research idea, involves multiple interdependent tasks, and requires the research path to be continuously adjusted in response to new literature, experimental results, and anomalous data.

Doing each Task well does not mean truly completing a Project.

Genuine project-level research requires AI to understand scientific objectives, proactively break down tasks, coordinate dependencies, execute continuously, and dynamically adjust its approach based on research feedback—ultimately driving the entire project toward a conclusion.

This has become a key capability that AI4S must develop in its next phase.

It is against this industry backdrop that ScienceClaw, the next-generation research-native agent developed by Zidong Taichu, completed a major upgrade.

The newly launched AutoProject project-level autonomous research engine aims to address the question of “how to conduct a research project more effectively,” driving AI beyond the execution of isolated tasks toward taking responsibility for complete scientific discovery.

The unit of work for AI in scientific research is shifting from Task to Project.

Real Scientific Research Must Be Able to Handle a Project

When taking on a genuine research Project, AI is not simply executing a single task. It must manage the complete research process, from project planning to the final conclusion.

AutoProject therefore does not merely extend AI’s task-execution chain. Instead, it builds project-level autonomous research capabilities around project planning, long-horizon execution, and evidence-based validation. Researchers need only provide a research idea or high-level objective, after which AI can autonomously complete the following process:

Project planning → SubTask decomposition → Long-horizon execution → Evidence validation → Dynamic remediation → Research output consolidation.

The system consists of three core capability layers:

  • Project2Task: Project-level planning — answers “How should this Project be conducted?”
  • TaskExecutor: Long-horizon autonomous execution — answers “Once the direction is set, how can the work continue moving forward?”
  • EviGraph: Evidence-driven validation — answers “Is the final conclusion sufficiently supported by evidence?”

Project2Task: Project-Level Planning and Autonomous Subtask Generation

When faced with a vague research objective, Project2Task integrates the research goals, literature evidence, resource constraints, and task dependencies to determine the research path autonomously. It breaks the work down into subtasks, plans the parallel and sequential relationships among them, and identifies opportunities to reuse assets across tasks, ultimately generating an executable and iterative project task network.

Take the intelligent identification of surface defects in steel strips as an example. Users need only enter a brief research objective, and AutoProject can use Project2Task to convert the vague requirement into a full-cycle task plan.

Video: https://mp.weixin.qq.com/s/Y6IKxQQjqq1Y6Fgwh94r9g

The system autonomously identifies tasks that can be performed in parallel and stages that must be executed sequentially, prioritizes multiple research paths, builds task dependencies, and discovers data and model assets that can be reused across tasks.

It also supports four project topologies—horizontal decomposition, vertical decomposition, horizontal-then-vertical decomposition, and vertical-then-horizontal decomposition. Based on the characteristics of each research topic, it generates a suitable research path and further translates high-level research objectives into specific, executable tasks.

From “humans define the Task and AI executes it” to “humans provide the objective and AI plans the path,” Project2Task enables AI to take on the top-level planning of research projects.

But planning is only the beginning. Once a clear project-level plan is in place, a reliable execution foundation is also needed.

TaskExecutor: Long-Horizon Autonomous Execution for Continuous Progress Through Dynamic Feedback

A research project is an iterative, evolving exploration loop—not a one-off delivery. An anomaly in any subtask can affect the upstream and downstream processes. Traditional AI based on single, one-time calls cannot adapt to the nonlinear, iterative nature of scientific research and struggles to support long-term studies.

AutoProject breaks away from the linear execution model of “one call, one response.” TaskExecutor is its implementation for research scenarios, providing project-level long-horizon autonomous execution.

Take YOLO modeling as an example. The project is divided into three major modules—data processing, modeling experiments, and iterative optimization—with each stage connected automatically.

After receiving the objective, the system autonomously studies papers, processes datasets, monitors training, handles anomalies, and automatically generates metrics, charts, and quantitative analyses.

Video: https://mp.weixin.qq.com/s/Y6IKxQQjqq1Y6Fgwh94r9g

TaskExecutor establishes a goal-driven autonomous research loop. It monitors the overall Project status, continuously captures experimental feedback, automatically investigates anomalies, and adjusts research hypotheses and model parameters.

When a task fails or produces anomalous results, the system can reassess and backtrack, reconstruct the task network, rerun experiments autonomously, and optimize the approach. It requires no frequent human intervention and continuously integrates fragmented experimental results into a complete research conclusion.

This means AI is no longer merely passively executing individual Tasks. It can make decisions based on the project’s status, iterate continuously, and drive the entire Project toward its objective.

Long-horizon closed-loop exploration also enables ScienceClaw to evolve from a tool-oriented research assistant into an agent with sustained autonomous research capabilities.

EviGraph: Evidence-Driven Validation for a Trustworthy Research Loop

As AI begins to execute a research project over an extended period, another question becomes increasingly important:

How reliable are the results it produces along the way?

The fundamental difference between scientific research and general-purpose generative content is that:

Scientific conclusions cannot be judged solely by whether the writing is fluent and the logic internally consistent. Every project-level conclusion must be grounded in a clearly defined research question, scientific hypotheses, multiple controlled experiments, and real data.

Zidong Taichu’s self-developed EviGraph evidence graph framework establishes a trustworthy research loop that covers the entire process, spans subtasks, and supports verification, traceability, and remediation.

Validation is no longer limited to individual subtasks. It also accounts for logical consistency across tasks:

The system continuously verifies whether the hypotheses align with the project’s top-level objectives, whether the experiments support the hypotheses, whether data from different subtasks corroborate one another, and whether the conclusions exceed the scope of the supporting evidence.

Once a logical deviation or data conflict is detected, the system can trace the evidence chain to identify the root cause, rerun experiments as needed, revise hypotheses, and adjust subtasks, enabling autonomous remediation.

Public testing has validated the effectiveness of this mechanism.

In the ARCBenchML evaluation, EviGraph achieved an overall score of 0.865, compared with 0.596 for the best baseline.

The Result Analysis metric, which measures how closely the conclusions align with experimental facts, increased from 0.442 to 0.794.

The Claim Support Rate for traceable evidence reached 0.38, a 40.74% increase over the best baseline of 0.27. Experimental Data Consistency (EDC) reached 0.88.

Once these three layers of capability are fully operational, the data, code, models, experimental records, and research methods generated during the research process can also be consolidated into reusable research assets.

In a study of proteins, these reusable assets included evidence summary tables, protein product files generated using specialized tools, research analysis reports, papers, and more.

Video: https://mp.weixin.qq.com/s/Y6IKxQQjqq1Y6Fgwh94r9g

It is worth noting that AutoProject’s autonomous research does not mean research without humans.

Even though the engine can complete the entire workflow—from planning and execution to validation and research output consolidation—researchers continue to play an irreplaceable central role.

Throughout the project, researchers can intervene at any time to review and assess the research paths, experimental hypotheses, and intermediate data produced by AI, promptly correct unreasonable directions, and contribute human domain intuition and scientific insight through ongoing interactive feedback. Humans and AI work together to complete the scientific exploration process.

Under this collaborative model, AI’s role also evolves from a research assistant that passively responds to requests into an autonomous driver of research projects.

AI4S Enters the Era of Project-Level Autonomous Research

As AI begins to work in terms of Projects, the boundaries of AI4S capabilities are changing as well.

Project-level autonomous research systems represented by AutoProject are driving a new paradigm shift across the industry:

AI’s object of work is moving from isolated Tasks to complete Projects, while autonomous research is evolving from “point capabilities” to “system capabilities.”

Achieving this transition is far from easy. So why was ScienceClaw the first to reach this stage?

The answer lies in the team and technology system behind it.

ScienceClaw was developed by Zidong Taichu, a company incubated by the Institute of Automation, Chinese Academy of Sciences.

It is a foundation-model company that has had “multimodality” embedded in its technical DNA since its founding.

As early as 2021, Zidong Taichu released the world’s first trillion-parameter tri-modal foundation model, bringing text, images, and speech into a unified semantic space.

After continuous iteration, today’s Zidong Taichu 4.0 has progressed from merely “understanding multimodality” to “multimodal reasoning,” with capabilities for autonomous planning, interleaved reasoning, and feedback-based correction.

In AI4S applications, the value of this technical path is clear:

Scientific research has never been a purely textual domain. Papers contain equations, charts, and experimental curves, while research projects also involve code, data, simulations, and real-world experiments.

Only a model capable of understanding these heterogeneous forms of information within a unified framework can truly serve as the brain of a research agent.

But foundation models are only the starting point.

Zidong Taichu has further extended model capabilities into real-world research workflows, building a complete technology stack spanning the model foundation, scientific knowledge, specialized tools, and agent execution.

This also sets ScienceClaw apart from ordinary research Agents.

ScienceClaw is not built on a traditional “master-slave Agent” architecture. Instead, it uses a layered-autonomy, dynamically collaborative multi-agent architecture designed for complex research tasks.

The system automatically generates task graphs around research objectives. Based on the domain, available tools, and execution status, it dispatches specialized Agents for disciplines, coding, search, data analysis, simulation, and other functions as needed, forming a research-agent cluster capable of dynamic network formation, collaborative execution, and feedback-driven replanning.

ScienceClaw currently covers fields including life sciences, materials science, chemistry, physics, and astronomy. It also integrates literature retrieval, data analysis, code execution, model computation, and simulation tools into a unified research-agent runtime.

Through cloud-based reasoning, sandbox execution, edge-side action, and coordination between virtual and physical environments, ScienceClaw is driving AI from “analyzing research information” toward “executing research tasks,” and from “completing a single Task” toward “continuously advancing an entire research Project.”

AutoProject represents a critical leap built on top of these capabilities.

On the surface, this is simply an upgrade in AI’s unit of work from Task to Project. At its core, however, it represents a leap in AI research capabilities from tool level to system level.

Along with this shift, the focus of competition in the AI4S sector is changing as well:

The industry is no longer competing solely on the performance of individual models or standalone agents, but on the overall capabilities of complete, end-to-end systems.

This is the core value of the latest ScienceClaw upgrade.

Zidong Taichu’s core advantage in AI4S lies not only in larger foundation models or higher-performing individual research Agents, but also in its complete research system, which integrates multimodal foundation models, research agents, specialized toolchains, autonomous execution frameworks, and experimental validation systems.

From Task to Project, ScienceClaw is driving AI4S into a new era of project-level autonomous research.

Paper links:

[1] https://arxiv.org/abs/2608.05225

[2] https://arxiv.org/abs/2608.04738