A New SOTA for Robot Brains
Overnight, the SOTA in robot brains changed hands.
This time, Dexmal’s DM0.5 has reached the summit of embodied AI—the “Everest” of robotics, RoboDojo.

This is RoboDojo we’re talking about…
It’s a devilishly difficult embodied AI benchmark, jointly launched by the Multimedia Lab at the University of Hong Kong, UC Berkeley, Tsinghua University, and nearly 20 other top academic institutions around the world.
How strict is the “quality control” at these elite universities?
Here’s some perspective: as of July 2026, the average success rate of the leading models on the simulation leaderboard was only 8.80%, while the real-robot leaderboard stood at just 12.8%.
Yes, they’re still struggling to break out of the single digits.
In short, it’s an extremely difficult and realistic benchmark—designed to expose models that excel only in isolated lab settings.
Now, DM0.5 has claimed first place:
- Overall score: 24.90
- Average success rate: 19.34%
There’s more. After taking a closer look at the table, I found that this model has one particularly outstanding strength—
Memory.
Take the Cover Blocks task, for example:
The robot first covers three randomly arranged red, green, and blue blocks from left to right. Once the colors are completely obscured, it must uncover them one by one in the order red, green, blue, based on its earlier observations.
DM0.5 achieved a 100% success rate across all three randomized trials.
The data reflects this very clearly: on the RoboDojo leaderboard, DM0.5 scored 47.74 points in the Memory dimension, far ahead of the competition…
That was undoubtedly a decisive factor in its climb to the top.

And that’s not all.
DM0.5 has also performed exceptionally well on other authoritative benchmarks.
- LIBERO: Overall success rate of 99.0%
- RoboTwin 2.0: Success rates of 93.6% and 93.3% in simple and complex scenarios, respectively
- VLA-Arena: Success rates of 89.0%, 53.6%, and 44.1% under different difficulty settings
- RoboChallenge Table30 v2: Overall score of 54.42 and a success rate of 43%
And this SOTA-level brain—
has already been open-sourced.
One Base Model, Six Categories of Scenarios
Leaderboard numbers are ultimately abstract. To understand what DM0.5 can really do, we need to look at its demos.
Dexmal is probably the foundation-model company with the richest collection of scenario demos at the moment.
Disturbance resistance, block assembly, cucumber peeling, parcel sorting… The range of demonstrations is dizzying.
But on closer inspection, these seemingly scattered demos all point to the same thing:
The generalization ability of the base model.
Let’s take a look at them one by one.
1. Human Demonstration
This is the demo I find the most interesting.
DM0.5 turns into Kakashi: show it something once, and it can replicate it at lightning speed.
To be honest, this capability is not exactly unfamiliar. Everyone is working in this direction.
After all, scenarios in the physical world are highly fragmented. No matter how capable a model is, there will always be long-tail tasks beyond the reach of its generalization abilities. Robustness drops sharply in these situations.
That makes it impossible to reach the ChatGPT moment for embodied AI.
The current industry consensus is not to expect pretraining to solve everything, but to introduce a Human-in-the-Loop during deployment. A human operator records a demonstration video, and the robot can immediately align itself with it—as if it had activated the Sharingan—and complete complex tasks accordingly.
Teaching a robot becomes more like training an apprentice.
That is exactly what Dexmal is doing.
And the “memory” mentioned at the beginning is key to this approach.
DM0.5 natively supports up to 60 seconds of memory, allowing the robot to deeply remember and continuously review every action sequence it has performed.
With this long-term memory, the model can directly overcome language barriers and understand video clips of human demonstrations.
2. Fine-Grained Manipulation
The scenario Dexmal chose here is block assembly.
I actually saw this demo at WAIC in Shanghai this past July. To be honest, it was far more impressive in person than in the video.
Dexmal teamed up with StepFun to assemble a Great Wall model in the exhibition area…
Yes, six robots spent 15 hours using 81,920 miniature blocks to build a 3.5-meter-long model of the Great Wall.

The crowds were intense too—it was practically impossible to squeeze through.
But this definitely wasn’t just a demo designed to generate “feel-good” reactions.
In the block assembly task, the smallest components are less than 1 centimeter wide. The robots must grasp and connect them within a scale of 0.1 to 1 millimeter. That level of precision surpasses the natural tremor range of human hands, which is approximately 0.3 to 1 millimeter.
At the same time, miniature blocks have manufacturing tolerances, so simply forcing them together can easily cause misalignment and jamming.
So if you watch the video closely, you’ll notice a pretty interesting detail.
This robot knows how to apply a short, sharp burst of force!
It first aligns the pieces, then gives them a gentle jolt and presses down with a slight shake to lock them together.

Of course, the ability to correct errors automatically is also crucial.
After all, it’s still a model. For high-precision tasks, you can’t realistically expect reliable Zero-shot performance every time. A closed loop is still needed to ensure stability.

Flexible objects require automatic error correction even more.
Take this cucumber-peeling demo. After each cut, the cucumber’s shape changes, so the model must continuously adjust its actions in a closed loop.
How deep to cut, how much force to apply, and how to peel it cleanly without damaging the flesh… Everything comes down to precision at the millimeter scale.

Think the cucumber has fulfilled its mission at this point?
Nope. Dexmal is still “whipping” the cucumber (just kidding).
But the focus is different this time: the robot uses a dexterous hand for cutting.
The model’s ability to handle high-degree-of-freedom manipulation through post-training is another demonstration of DM0.5’s generalization capabilities.

3. Long-Horizon Stability and High Throughput
From here on, the demos become closely integrated with real-world scenarios.
At this year’s WRC in Beijing’s Yizhuang district, Dexmal brought a parcel-sorting line to the venue to tackle the biggest pain point in logistics transfer stations—
the singulation of individual items.
The conveyor belt keeps moving, while every parcel differs in size, material, and orientation. Traditional mechanical solutions simply cannot handle this type of task.
DM0.5 relies on neither predefined scripts nor hand-crafted routines. Using real-time visual recognition and decision-making alone, it processes one parcel every three seconds on average, with an accuracy rate above 99%.

I also saw this demo up close at the venue last week.
To be honest, when it comes to pure entertainment value, I personally think industrial scenarios are the best. They’re incredibly satisfying—like watching cleaning livestreams.
At least they’re easy to watch. The robot doesn’t move painfully slowly or freeze up while picking up a pillow…
Highly recommended viewing for anyone with ADHD (just kidding).
4. Disturbance Resistance
The real world is not a laboratory. The camera’s viewpoint can change abruptly at any time, and humans can also suddenly “drop in” and cause trouble.
DM0.5’s solution is a hierarchical dual-system architecture, namely the industry-standard System 1 and System 2 architecture.
System 2 is the deliberative, higher-level brain. It handles understanding, refusal, and global anticipation, with Zero-shot performance as its goal.
System 1 is the action-expert network. Like human reflexes, it handles the trivial details involved in long-horizon, complex tasks.
With this architecture, even when the external viewpoint changes dramatically, DM0.5’s general-purpose picking remains highly robust.
Even when a human forcibly interrupts a task, the robot can accurately understand the current state of the environment and adaptively correct its subsequent actions.
Six categories of scenarios—from millimeter-scale manipulation to conveyor-belt sorting, from rigid objects to flexible objects, and from execution to imitation.
It is rare for a single model to perform so well across scenarios that differ this greatly.
Everyone is racing toward millions of hours of data, but focusing solely on data volume may not align with first principles.
If the base model lacks strong generalization capabilities, and a separate model has to be trained for every scenario, then no matter how many demos you have, they are nothing more than a pile of scattered pearls.
A model with narrow strengths cannot achieve Scaling Up.
That is also what I find most interesting about DM0.5: the capabilities of embodied models themselves are beginning to be recognized, and their ceiling can be remarkably high.
Comprehensive Upgrades to Data, Architecture, and Efficiency
So how exactly did Dexmal achieve all this?
The team has already shared its experience openly and in full.
The answer has three layers, corresponding to the three core variables that embodied AI cannot avoid: data, architecture, and latency.
Let’s start with the data.
Data determines the upper limit of intelligence, making it the single most important component in the entire embodied-AI engineering stack.
DM0.5’s data assets are primarily divided into three categories:
- Real robots: 50,000 hours of high-precision manipulation data, covering more than 100 types of actions and enabling second-level alignment between fine-grained instructions and actions.
- Ego: 100,000 hours of scenario video, supporting millimeter-level, high-precision 3D Landmark generation for accurate annotation.

- Simulation: Modeling data across 1 million square meters of space to build complex indoor environments, using high-precision reconstruction to address the Sim2Real Gap.

It feels like the team places a strong emphasis on quality rather than simply collecting data impulsively and indiscriminately.
In terms of data volume, they use an Omni-fusion approach, which is also one of the industry’s mainstream solutions and minimizes risk.
Now let’s look at the architecture.
In a nutshell, DM0.5 features three major innovations designed to solve three key problems.
1. How Can It Remember?
We touched on this earlier. To lower the barrier to deploying DM0.5, Dexmal introduced native long-term memory.
More specifically, the team carried out extensive work at the context abstraction layer. Through efficient information-compression techniques, they enabled a 4B-parameter VLM to natively learn and remember historical information during pretraining.

2. How Can It Understand What It Sees?
Dexmal has reached the following conclusion:
A VLA without language capabilities is just a dataset parrot.
After all, language is by far the most convenient modality for distilling human reasoning abilities. Having a language layer is a significant advantage.
This led to the emergence of embodied chain-of-thought tasks.
To this end, DM0.5 designed embodied chain-of-thought tasks and used 11 reasoning tasks to drive the model to jointly model instructions, its own embodiment, and the environment—giving action generation semantic understanding and the ability to anticipate the world.
During this process, the team also created “counterfactual tasks,” deliberately providing incorrect instructions to force the model to genuinely “understand” what it sees rather than mechanically match patterns.
3. How Can It Align Its Actions?
Traditional action supervision is point-based. The deviation between the predicted trajectory and the ground-truth trajectory accumulates continuously, eventually leading to failure.
Dexmal uses alignment-based action supervision and efficiently computes the optimal matching through constrained dynamic programming, resolving inconsistencies between action trajectories in a single step.

The combination of data and architecture has significantly raised DM0.5’s intelligence ceiling.
But to make the system truly useful, efficiency is just as important.
Embodied models ultimately need to run in real time on edge devices. If inference isn’t fast enough, all those parameters are just a useless pile of paper.
To address this, Dexmal carried out a series of optimizations, which I won’t go into here. You can read the official technical report directly.
The final results are as follows:
- Core model latency: 534.04 ms → 57.49 ms
- Cumulative speedup: 9.29×
- Latency reduction: 89.2%
- API latency: p50 at 67.75 ms and p90 at 70.01 ms
- LIBERO success rate: 98.45% → 98.40%
That’s a full order-of-magnitude speedup…
And with virtually no loss in accuracy.
You could say they’ve enabled the entire VLA inference system to undergo a major evolution.

Reaching the Summit Is for the Descent
Reaching the summit is certainly impressive, but to be honest, it may not be the thing most worth paying attention to.
In embodied AI, one day in the heavens is a year on Earth: the championship changes hands constantly. A single scorecard can quickly become outdated.
What truly leaves a mark on the development of embodied AI is what you are willing to leave behind after reaching the top.
That may also be why DM0.5 chose to be fully open-sourced.
It’s not just the model weights. DM0.5’s training framework and downstream task workflows… are being released progressively on GitHub and Hugging Face, allowing developers to reproduce, modify, and extend the system.
For example, the demo below was trained by Dexmal on top of DM0.5, combined with an open-source robotic-arm kit and supervised fine-tuning.

In the LLM field, open source is no longer anything new. From DeepSeek at the beginning to GLM, Kimi, and MiniMax today… everyone has grown accustomed to seeing models released openly.
But in the embodied-AI industry, I think this means something different.
The fact is that physical AI is still at a very early and immature stage.
So we need more transparent discussion, more realistic evaluations, and more publicly available work…
If Google had not released the Transformer back then, OpenAI might never have found its way to ChatGPT.
The popularity of RoboDojo has pushed evaluation away from “single-demo showcases” toward standardized, comprehensive examinations, making it possible for embodied AI to be measured, reproduced, and challenged under a unified framework.
DM0.5’s ascent to the top—and its open-source release—further delivers on the value of this benchmark. It shows everyone the methodology behind reaching the summit, while making that methodology available to take away and reuse.
This is Dexmal’s answer to the embodied-AI race underway today.
Climbing Mount Everest has never been about staying at the summit.
The next step is to descend—
to empower the ecosystem through open source and bring the embodied brains of the first tier onto production lines across industries.
GitHub:https://github.com/dexmal/opendm
Hugging Face:https://huggingface.co/Dexmal/DM05
Tech Blog:https://www.dexmal.com/blog/dm0.5