Skip to content
Main Site News Console

WRC 2026 | ShengShu Technology Unveils Latest Achievements and Proposes a Five-Level Roadmap for General-Purpose World Models

· 量子位
国内AI

On the afternoon of August 19, the World Robot Conference subforum “The Evolution of Embodied Intelligence Foundation Models: From Technical Hierarchies to Industrial Deployment” was held at the Beiren Yichuang International Convention and Exhibition Center. Zhu Jun, founder and chief scientist of ShengShu Technology and an ACM/IEEE/AAAI Fellow, delivered a keynote speech, unveiled the team’s latest research成果, and systematically outlined a five-level development roadmap for general-purpose world models.

Zhu Jun said: “From the perspective of foundation model development, what we hope to build is not merely a specialized model designed for a particular task or scenario, but a general-purpose foundation model that can understand the world, predict the future, and take action.”

Defining General-Purpose World Models from First Principles

World models have expanded into areas such as video generation, environment simulation, robot decision-making, and action control. However, the industry still lacks a consensus on what qualifies as a truly general-purpose world model.

When people learn to ride a bicycle or drive a car, their movements gradually become stable and precise after an initially uncoordinated phase. One important reason is that, through continuous interaction with the external world, the brain develops an “internal model” of how the world works.

A general-purpose world model likewise requires three interconnected capabilities:

  • Understanding the world: interpreting the environment and determining its current state;
  • Predicting the future: forecasting what may happen next and what outcomes different actions may produce;
  • Taking action: affecting a digital or physical environment and using real-world feedback to refine its understanding and predictions.

“A general-purpose world model is not simply a generator, simulator, robot action model, or policy model. Nor is it a collection of disconnected capabilities or a simple linear combination of them. Rather, it is a closed-loop feedback system in which all three capabilities are tightly coupled,” Zhu Jun emphasized. Actions are not merely model outputs; they also alter the environment, generate new information, and feed into the next round of understanding, prediction, and decision-making.

Data, Architecture, and Compute: The Three Essential Elements of General-Purpose World Models

Building a general-purpose world model still depends on the three fundamental elements of large models: data, architecture, and compute.

On the data side, a general-purpose world model needs a multilayer data pyramid that progressively moves from observation to action: beginning with massive volumes of online video, then expanding to domain-specific video, first-person human video, human demonstrations with action records, and real-world robot interaction data.

The closer the data is to the base of the pyramid, the larger its scale and broader its coverage. Toward the top, the data becomes scarcer and more expensive, but it can establish more direct connections between tasks, actions, and physical outcomes. Synthetic data can be used across every layer of the pyramid. At the same time, “imperfect data” is also highly valuable: failed attempts, corrective actions, and recovery processes all contain useful learning signals that can help robots learn how to recover from failure.

At the architectural level, a general-purpose world model needs to process different modalities—including images, video, language, and robot actions—in a unified manner. MoT (Mixture-of-Transformers) assigns dedicated parameters to different modalities while using shared attention to enable cross-modal information exchange. This allows environment understanding, world-state prediction, and action generation to work together within a single model. ShengShu Technology’s world action model Motubrain is built on this architecture, unifying understanding, prediction, and action within a single modeling framework. It overcomes the fragmentation of capabilities found in traditional modular approaches while improving the efficiency of heterogeneous data utilization and cross-task transfer.

Compute affects both large-scale pretraining and real-time deployment. Offline video generation can tolerate some waiting time, but real-time interaction and robot control must complete prediction and decision-making before the environment changes. As a result, training infrastructure, inference acceleration, model distillation, and efficient attention mechanisms are all essential to bringing general-purpose world models into practical use.

How Do General-Purpose World Models Evolve? A Five-Level Roadmap

Starting from the goal of building a general-purpose foundation model, general-purpose world models can be divided into five progressively evolving levels. These are not isolated product directions, but stages in a process through which models continuously deepen their understanding of the world, strengthen their interaction with it, and increase their autonomy. Over the past several years, ShengShu Technology’s research and deployment efforts have already covered the first three levels.

L1: World Generation

The first step is to enable models to learn coherent processes of world evolution. Video is the most representative medium at this stage. It helps models learn about objects, motion, spatial relationships, and temporal changes, laying the foundation for further understanding and prediction of the world.

In 2024, ShengShu Technology launched the video generation foundation model Vidu, and through continuous improvements in video quality, temporal consistency, and world-dynamics modeling, completed its initial implementation of the L1 stage.

L2: Interactive World

Building on generation capabilities, models begin to receive real-time inputs and continuously modify subsequent content based on language, speech, or control signals, moving from “one-time generation” to “continuous interaction.”

Vidu S1, released in July 2026, advanced this capability into the real-time interaction stage. Users can intervene in the generation process in real time through voice and continuously influence what happens next, enabling the model to dynamically evolve the world based on external input.

L3: Actionable World

The next step is for models to move from the digital world into the physical world. At this stage, environment understanding, future prediction, and action generation must be genuinely unified. Models must not only determine what is happening in the world, but also directly generate executable robot actions and continuously refine them based on real-world feedback.

Motus and Motubrain mark ShengShu Technology’s expansion from video generation into physical action. Motus was released and fully open-sourced in December 2025; Motubrain, released in April 2026, further unifies environment understanding, world-state prediction, and action-trajectory generation within a single model, advancing general-purpose world models from “visual simulation” toward “physical decision-making.”

Compared with Motus, Motubrain delivers approximately 10 times faster inference. When adapting it to a new robot platform, only 50 to 100 human demonstration samples are required. In the RoboTwin 2.0 benchmark, it achieved a score of 96.1, ranking first. To date, Motubrain has been validated on nearly ten robot platforms, including those from Galbot, Astribot, and X Square Robot, demonstrating its ability to generalize across long-horizon tasks, multi-task scenarios, and different robot embodiments.

L4: Autonomous World Agent

After gaining the ability to act in the real world, models must advance toward autonomous decision-making. They need to proactively perceive their surroundings, break down tasks, explore unknown states, and continuously plan actions around long-term goals.

L5: World Orchestrator

At a higher level, models will no longer control only a single agent. Instead, they will be able to coordinate robots, digital agents, humans, tools, and other resources to carry out more complex task allocation and multi-agent collaboration.

From L1 to L5, capabilities advance step by step. Without learning and predicting the laws governing how the world evolves, it is difficult to achieve stable, continuous interaction. Without environmental feedback from real-world action, it is equally difficult to develop autonomous learning, long-term planning, and complex coordination.

Six Key Challenges Remain on the Path to L4 and L5

L1 through L3 have already given rise to relatively concrete technical implementations, but a gap remains before higher-level world intelligence can be achieved. Video models still need to further improve generation quality, controllability, and spatiotemporal consistency. World action models, meanwhile, need to enhance their stability, generalization, and execution efficiency in more complex and open-ended real-world environments.

Advancing toward L4 and L5 will require breakthroughs in six key areas:

  • Establishing a joint evaluation system covering understanding, prediction, action, and transfer capabilities;
  • Learning real-world physical laws, including contact, friction, force, and the consequences of intervention;
  • Developing persistent, revisable memory;
  • Achieving online learning and self-evolution through real-world interaction;
  • Implementing an efficient closed loop that meets real-world latency and resource constraints;
  • Further strengthening safety and controllability.

With this direction in mind, ShengShu Technology will continue to strengthen the three levels of world generation, interaction with the world, and action in the world, while further developing capabilities such as goal formation, active exploration, continual learning, and long-term memory. These efforts will lay the foundation for higher-level autonomous world agents. The longer-term goal is to build a general-purpose foundation model that can continuously understand the world, predict the consequences of different actions, act autonomously within clearly defined constraints, and continuously learn and evolve from real-world feedback.

When understanding, prediction, and action truly form a closed loop, world models will no longer be merely models for generating content or robot policies for executing tasks. Instead, they will become an important foundation for connecting the digital and physical worlds and advancing toward general-purpose embodied intelligence.

Blog link: https://www.shengshu.com/zh/general-world-model/

This article was provided by ShengShu Technology and reprinted by QbitAI with authorization. The views expressed are those of the original author.