Skip to content
Main Site News Console

WRC 2026 | Natively Multimodal World Models: From Simulated to Interactive Worlds

· 量子位
国内AI

On August 19, during the World Robot Conference (WRC 2026), the “Physical AI Leaders Forum and Launch Ceremony for the 2026 China AI 100 List by Caijing,” hosted by Cross-Dimensional Intelligence and co-organized by Caijing magazine, was held in Beijing. With the theme “Physical AI in Progress,” the forum brought together leading scholars, company founders, and frontline technical leaders in embodied intelligence to explore the technological evolution and industrial pathways of general-purpose physical intelligence.

As a representative company in native omni-modal world models, Tao Mei, founder and CEO of HiDream.ai, delivered a keynote speech titled “Native Omni-Modal World Models: From ‘Simulating the World’ to ‘Interacting with the World.’” Starting from the technological evolution of AI foundation models, he systematically explained the underlying logic behind the transition from multimodal to native omni-modal models and from simulated worlds to interactive worlds, offering new technical perspectives and potential pathways for building the foundational capabilities of Physical AI.

01 High IQ Does Not Mean All-Powerful: AI’s Transition from “Understanding the World” to “Changing the World”

In his speech, Dr. Tao Mei noted that AI foundation models have advanced at an astonishing pace over the past few years. From large language models to multimodal models and then to embodied intelligence, three technological paths that once evolved independently are now converging rapidly. Many have called 2026 “the first year of world models.”

He shared a set of figures: at the beginning of this year, the IQ level of top foundation models was approximately 130, comparable to that of gifted students in the special program for young students at the University of Science and Technology of China. Today, the IQ of the latest foundation models is approaching 140. However, Dr. Tao emphasized a critical point: “A high IQ does not mean being all-powerful.” Just as good grades do not necessarily indicate well-rounded abilities in a child, a model that performs exceptionally well on tests is not necessarily capable of completing tasks over extended periods in the real physical world consistently, reliably, and robustly. This is precisely why Physical AI is so important.

Large language models excel at understanding the world, compressing knowledge, and reasoning. Multimodal models excel at generating and predicting the world, while embodied interaction models excel at interacting with the world. A true world model, however, must simultaneously provide a complete representation of the world’s states, reason about physical laws and causal relationships, and reconstruct the world. The three paths that previously evolved independently—from text to multimodality to embodiment—are now converging rapidly.

02 From DiT to UiT: A Native Omni-Modal Architecture Breaks Through Industry Limits

Built on its self-developed UiT (Unified Transformer) architecture, HiDream.ai recently officially released HiDream-O1-World, a native omni-modal interactive world model. Upon its first evaluation, the model topped the Navi sub-ranking of the authoritative WBench leaderboard.

HiDream-O1-World supports multimodal inputs, including text, images, and interactive control. It offers three core capabilities—navigation, editing, and interaction—while also supporting the creation of scenes featuring diverse subjects and styles. The model can both highly faithfully reproduce real-world spaces and generate stylized worlds in formats such as 2D animation and AAA game rendering.

At the technical level, HiDream-O1-World’s core breakthroughs lie in long-horizon spatiotemporal consistency and physical consistency. These capabilities allow the model to “remember” the structure of explored scenes during long-sequence interactions and maintain a high degree of physical plausibility during complex interactions.

HiDream-O1-World is opening up new possibilities for industrial applications. In AI-powered interactive films and games, users become active participants in the narrative. In embodied intelligence simulation, the model can provide a high-fidelity virtual foundation for training. In 3D scene generation, creators can efficiently generate structurally complete 3D environments, significantly shortening the iteration cycle from concept to finished product.

With this as its starting point, HiDream.ai is continuing to expand the boundaries of world models and gradually refining a model portfolio spanning image models, video models, and interactive world models.

03 Data · Models · Agents · Embodiment: Closing the Loop with the Physical World

Discussing the capability foundation of embodied intelligence, Dr. Tao Mei summarized the key relationship as follows: “The development of world models and embodied intelligence requires connecting data, world models, agents, and embodied platforms.” Data builds cognition; world models are responsible for understanding the state of the world and reasoning about physical laws; agents make decisions and plans based on predictions; and embodied platforms carry out actions in the real world. Feedback from the real environment then flows back into the system as new training data, creating a closed-loop flywheel of “data—cognition—action—feedback.” Dr. Tao Mei stated that world models serve as the cognitive hub of this flywheel. Without high-quality world models, simulation cannot adequately approximate real physical environments, and the data flywheel for embodied intelligence cannot truly get moving.

At the foundation is the data layer, comprising three indispensable categories: “prior data from the internet, physical trial-and-error data from simulations, and real interactive data from closed-loop actions.” Feedback from actions in real-world environments flows back into the system as new training data, continuously accelerating the flywheel. Taking internet-based prior data as an example, HiDream.ai has accumulated nearly 10 million hours of video data.

The middle layer is the world model, which serves as the core of embodied intelligence. It is responsible for reasoning about the laws of the world and answering the questions, “What is the world?” and “What will happen next?” At the top layer are agents and embodied systems, which make autonomous decisions based on the predictions of the world model. Embodied platforms then execute physical actions, completing the full loop from cognition to action.

Dr. Tao Mei placed particular emphasis on the value of simulation data for Physical AI. He cited HiDream.ai’s collaboration with Noitom as an example: by using a native omni-modal foundation model to augment real-world human motion data, the team can generate large volumes of videos that comply with physical constraints. The same action remains consistent across different backgrounds, settings, and skin tones. This data augmentation capability can significantly reduce the cost and burden of collecting real-world data, allowing robots to learn first in simulated environments and then train in real-world settings.

Underlying these efforts is HiDream.ai’s consistent view of the future direction of AI: foundation models enable AI to understand the world, while agents and embodied systems enable AI to act in the world. Native omni-modal world models connect cognition, action, and feedback into a unified whole, giving AI the ability to continuously learn and evolve in the physical world.

From simulating the world to understanding it, and from reasoning about the world to interacting with it, HiDream.ai will continue to focus on native omni-modal world models. Driven by both technological innovation and ecosystem collaboration, the company aims to serve digital intelligence while accelerating the training and deployment of embodied intelligence in real physical environments—helping Physical AI move from technological exploration into a new stage of large-scale commercialization.