Not a Single Fresh Twist?
Hengyu, reporting from Aofeisi Temple
During WRC, the most eye-catching attractions at the robot booths across the exhibition hall were all kinds of novel movements.
A robot dancing, shaking hands, or unlocking a complex task that people had never seen before—as long as it was intuitive and novel enough, visitors would pull out their phones to record it.
But once robots are truly deployed in factories, shopping malls, or ordinary homes, practical concerns return to the most fundamental level: in real-world scenarios, what matters more is whether the robot can perform the task faster and more reliably next time, while minimizing human intervention and assistance.
That is precisely why Qianxun Intelligence’s approach at the 2026 World Robot Conference stood out.
Rather than rushing to unveil a fresher demo scenario, they had Moz1 bring the “tidying up the living room” demo showcased at WAIC more than a month earlier to WRC for an encore.
Hmm… reheating leftovers?

So, at the exhibition hall, we found a Qianxun Intelligence employee and asked: Was this really just the same demo on display? Not a single fresh twist?
They smiled and replied, “Don’t just look at the surface! Although it’s been less than 30 days, this demo has already made significant progress:
- The success rate for autonomously reasoning through putting a can of Coke in the refrigerator rose from approximately 80% to more than 99%;
- The success rate for putting a bowl into the dishwasher rose from approximately 90% to more than 99%;
- The time required to confirm navigation targets fell by nearly 50%.
In addition, when conditions such as color, type, and position changed, the success rate for the litter-picking task still reached as high as 85%.

These changes occurred at different stages of the long-horizon compound task of “tidying up the living room,” rather than in a single pick-and-place action or a fixed workstation operation.
So the two appearances of the same demo were actually two “evolution snapshots” of the same embodied system.
At first glance, it may seem that the robot is still doing the same old thing. But beneath the surface lies an iceberg: above the waterline is improved task completion; below the waterline is progress across the full technology stack, encompassing data, training, agents, navigation, hardware, and infrastructure.
How did this rapid evolution over 30 days happen?
We need to dive beneath the surface.
A Quantifiable Leap in the Demo over 30 Days
Understanding the changes over these 30 days requires a term called “time density.”
This is the first academic concept Qianxun Intelligence has proposed in the field of embodied intelligence. “Time density” focuses on how quickly a system can absorb problems, complete training and validation, and redeploy improvements to a real robot.
It measures not the peak performance of a particular demo, but the speed at which the system evolves.

In fact, Qianxun Intelligence’s success rates were already fairly high 30 days ago: during WAIC, the success rate for putting a can of Coke in the refrigerator was close to 80%, while the success rate for putting a bowl in the dishwasher was close to 90%.
Those were impressive results. If the goal were simply to complete a trade-show demo, that level of performance would already have been enough to produce a solid on-site showing.
But achieving 90% on a long-horizon task is entirely different from achieving 90% at a fixed workstation.
A single-step task that fails once usually means that a particular action was not completed. In a long-horizon task, however, a small deviation can change the state encountered in the next step, allowing the error to continue propagating downstream.
The series of everyday actions involved in having the robot “tidy up the living room” is actually quite complex.
The robot must understand ambiguous instructions, first scan the environment, determine which objects need to be put away, and remember which steps remain unfinished. It then has to move across spaces, grasp objects, open doors, place items, close doors, and continuously confirm along the way that it is doing things correctly.
Moreover, in long-horizon tasks, changes in an object’s pose, shifts in the robot’s position, background interference, and even an insufficient door-opening angle can all disrupt the rhythm of subsequent actions.
Therefore, a high success rate in a demo is only the opening chapter. It does not mean the robot has mastered the task flawlessly.
Let’s return to the WRC floor.
The four task groups—refrigerator operation, dishwasher operation, picking up litter, and putting away dolls—each exposed different problems in long-horizon tasks: whether consecutive operations can be linked reliably, whether the task can continue when the robot’s position and the objects’ states change, and whether targets can still be identified accurately after the environment changes.

After watching a few more rounds of demonstrations, it became clear that the difficulty was constantly increasing on the exhibition floor.
Visitors walking through the area, occlusions, and changes in object positions could all temporarily disrupt the route. The robot had to continue avoiding obstacles, approach the target again, and reconnect navigation with the subsequent operation.
It was plainly obvious that exhibition halls are very good at creating trouble.
But for embodied robots, those troubles happen to be valuable.
(To all our fellow exhibitors: enjoy it while it lasts!)
In addition, the stability of on-site performance is difficult to achieve through parameter tuning of a single model alone.
Only when every failure can be identified, fed back, reproduced, trained on, and redeployed can a closed data loop be formed to support rapid iteration within 30 days.
The Full-Stack Iceberg Beneath the Waterline: How to Run a Long-Horizon Task
Breaking down the complete execution process for “tidying up the living room” is essentially the same as breaking down Qianxun Intelligence’s full-stack technology system.
For a robot, the instruction “tidy up the living room” must be broken down into a series of questions:
What is in the room? Which items need to be put away? What should be done first and what should come next? Where should the robot go to complete the operation? If something fails midway, what should it do next?
Once the user issues the goal of “tidying up the living room,” the perception system first scans the entire environment and builds an understanding of the surroundings. The Agent preserves context, remembers the long-term goal, and breaks the larger task down into subtasks such as “put away the Coke first, then wash the dishes.” The navigation system takes the robot to a suitable operating position, while the base model Spirit v1.6 generates actions based on language, vision, and the robot’s proprioceptive state.
After the task is completed, the system must also generate a retrospective report and feed the data back into the system so that performance can improve next time.
As we can see, behind a complete task is the coordinated work of perception, data, models, agents, navigation, the robot body, and infrastructure.

△ Qianxun Intelligence technology architecture
Every link in this chain is interconnected and indispensable, enabling the robot to continuously learn from each task.
Let’s start with the data pipeline—the “fuel” for system evolution.
To make the robot better 30 days from now, the first problem to solve is: Why did it fail this time? What exactly should be changed in the next training cycle?
For example, a failed attempt to pick up litter does not mean the task is over. What really matters is whether the system knows why it failed: Did it fail to recognize the litter, was the grasping position incorrect, or was there a problem with action planning?
The answer lies in the data.
Qianxun Intelligence’s data platform brings together web videos, first-person-view data, uDAS data, teleoperation data, and real-robot task data. It covers data cleaning, annotation, quality inspection, training-task management, model evaluation, and result tracking.
At present, Qianxun Intelligence’s self-developed data-collection equipment has reached its seventh generation. Data-collection costs have fallen to one-tenth those of traditional teleoperation, while data usability has risen from 30% to 95%. Together with more than 300,000 collection sites across China, the system continuously supplies “fuel.”

△ Qianxun Intelligence’s wearable data-collection devices collect data simultaneously in many parts of China
For embodied intelligence, however, the most valuable data is not necessarily the “cleanest” data.
This brings us to Qianxun Intelligence’s original “dirty data” training philosophy and data pyramid system.

△ Qianxun Intelligence data pyramid
In the real world, lighting changes, objects do not remain in fixed positions forever, people may suddenly walk by, and the environment is constantly shifting.
If the robot sees only orderly, standardized scenes during training, it may actually be more susceptible to interference from irrelevant information when facing the real world.
Qianxun Intelligence therefore emphasizes “dirty data” training, directly incorporating the messy variations of the real world into the model’s generalization capabilities.
This also needs to be considered together with the post-training process.
The post-training mechanism and base model are the foundation that determines the upper limit of the robot’s action capabilities.
The core challenge during post-training is to enable the model to genuinely understand experience and know “what is worth paying attention to.”
Throughout the task, the robot must keep its eyes and ears open in every direction.
Take “putting a bowl into the dishwasher” as an example. This task contains at least 10 consecutive stages.
At the beginning of the task, global vision helps the robot quickly determine its own location and the location of the dishwasher. Once the operation actually begins, the area of focus gradually narrows. At this point, what matters more is the position of the bowl, the position of the handle, and the position of the dish rack.

Through post-training, the team taught the model to remain stable in response to irrelevant changes, such as walls and lighting, while remaining sensitive to changes that truly affect the action, such as the bowl’s position and the robot’s stance.
Data and post-training ultimately need to translate into model capabilities, and the foundation supporting all of this is the Spirit v1.6 base model.
According to Qianxun Intelligence, the model uses a deeply integrated architecture combining a VLA with a world model.
Traditional robots often use a step-by-step “perception–planning–control” process—like looking at a map first, planning a route second, and stepping on the accelerator last. This is slower to respond and more prone to breaks between stages.
By contrast, Spirit v1.6 incorporates visual perception, language understanding, action generation, and world-state prediction into unified training. It can perceive, make decisions, and adjust simultaneously—like an experienced driver whose hands and feet instinctively coordinate when an unexpected situation arises.
In the RoboArena international evaluation jointly launched by UC Berkeley, Stanford, and NVIDIA, the model once ranked first worldwide, surpassing models such as NVIDIA’s DreamZero. It was the first Chinese embodied model to top the leaderboard.

At the Agent decision-making layer, the “brain” must prevent the robot from developing a truly “goldfish-like” memory during long-horizon tasks.
More specifically, the biggest fear is that the robot will simply stop where it is halfway through a task, forgetting why it is standing there.
For example, after a failed attempt to pick up litter, the action module may issue a “stop” signal. The planner mistakenly assumes that the current step has been completed and moves directly to the “put away the litter” task. A single-step misjudgment thus propagates along the task chain.

To prevent this chain reaction, the Agent acts like a foreman carrying a notebook and must continuously manage the task state.
It must not only preserve the overall goal of “tidying up the living room” and the progress already made, but also decide based on the current environment whether to continue, retry, skip, or replan.
For example, after a failed litter pickup, should it try again or switch to another target? If someone blocks the route, should it wait or go around?
Only when the Agent keeps the state continuously aligned can the robot avoid going “blank” halfway through the task.
An unexpected test at WAIC demonstrated the capabilities of this “foreman.”
When Qihoo 360 Group’s Zhou Hongyi and his delegation visited the booth, the living-room demo area was packed with people. The static living room instantly became a highly dynamic, heavily occluded environment.
The Agent monitored task progress throughout: when navigation was blocked, it determined whether to wait or take a detour; when someone approached the operating area, it paused and resumed only after confirming that it was safe; when the target was occluded, it initiated perception again.
In the end, the robot navigated through the crowd, avoided obstacles, and completed a sequence of operations without interrupting the task.
The coordination between navigation and operation reflects an upgraded definition of what it means to “arrive.”
The robot stopping in front of the dishwasher does not mean it has truly arrived. It must stand in a position from which the VLA can open the door, pull out the dish rack, and place the bowl. Only then has it completed this stage of the task.
There is a classic story from the development process.
During the early debugging phase, the robot was already able to walk smoothly to the dishwasher.
But on one occasion, it stopped too close to the refrigerator. Judging solely by navigation metrics, it had “arrived.” Yet as soon as the policy model took over, the problem became obvious: this position was unfavorable for opening the dishwasher, requiring substantial additional pose compensation.

The team therefore redefined what “arrival” meant.
The robot could no longer use reaching a specific coordinate as the standard. What it truly needed to reach was a position from which the subsequent action could be carried out—in other words, an “operable pose.”
Standing in the right place is not enough; the robot also needs to respond quickly.
Another optimization focused on the speed at which navigation begins. Through coordinated adjustments to the model, algorithms, and engineering pipeline, the median time from receiving a task to generating the first navigation target was reduced by nearly 50%.
All of the above concerns only the software layer. Hardware and infrastructure should not be forgotten.
During WAIC, the team encountered problems including servo-motor abnormalities, power failures, and loose interfaces.
After running for more than ten minutes, a single interface failed at the final step, causing the entire task to end in failure. In a long-horizon task, encountering this kind of problem is no exaggeration—it genuinely feels to the team as if “the sky is falling.”
After WAIC, the hardware team therefore changed its strategy, replacing “fix things wherever they break” with “bringing design quality forward.” They worked to intercept problems at the drawing stage, continued improving the servo and power modules, and embedded design reviews and assembly checks into the standard process.
At the same time, the actions ultimately generated by the model still need to be executed by Moz1’s 26-degree-of-freedom robot body, equipped with self-developed integrated force-controlled joints. How quickly a new model can run on the real robot also depends on underlying platform capabilities such as device management, inference deployment, and log feedback.
This invisible infrastructure further shortens the distance between training and the real robot, allowing bad cases encountered on-site to enter the next iteration more quickly.
After Reaching a 100% Success Rate, What Is the Point of Continuing to “Find Faults”?
Drawing on its existing model capabilities, Qianxun Intelligence achieved zero-shot transfer directly in the new WRC scenarios. The success rates for several core tasks—including refrigerator operation, doll storage, bowl and Coke grasping, and dishwasher operation—reached 100% in on-site testing, demonstrating strong cross-scenario generalization.
But after reaching a 100% success rate, the team instead began actively “finding faults,” moving the starting position, changing object poses, replacing the background, and so on.
Because they found that achieving 100% in a single scenario does not mean the stability boundary is already broad enough.
The details exposed during real-robot debugging further illustrate the complexity of long-horizon tasks.
After one failed litter-grasping attempt, the action module issued a stop signal. The planner mistakenly assumed that the current step had been completed and moved directly to the task of putting away the litter. A single-step misjudgment continued to propagate through the long-horizon task.
Therefore, long-horizon tasks cannot be judged solely by whether an action has ended. The system must continually confirm: “Was this step actually completed successfully?” Otherwise, a small misjudgment will keep rolling forward.
Network conditions at the exhibition hall also caused considerable trouble.
The venue lacked a stable wired network, and cloud inference was subject to unpredictable long-tail latency.

To address this, the team carried out targeted optimizations: removing redundant decision logic, converting some visual judgment tasks to parallel execution, and overlapping inference computation with the robot’s other actions as much as possible to mitigate the impact of network latency.
Once large numbers of visitors entered the exhibition hall, the relatively static environment became a highly dynamic, heavily occluded space. Pedestrians could temporarily block the target, while planned routes could suddenly become unusable. Navigation had to continuously avoid obstacles and re-approach the target, while also keeping task states aligned with the Agent and VLA.
From this perspective, the exhibition floor became a ready-made stress-testing environment—one that could, precisely because of these conditions, demonstrate the strength of the full technology stack from another angle.
At the same time, the extreme stress testing at the exhibition hall served as a rehearsal for stable deployment in industrial environments. Although the two settings differ, both test whether a robot can operate for extended periods with stability after leaving the laboratory.
Qianxun Intelligence’s system capabilities have already begun expanding beyond exhibition demos into real-world industrial scenarios and others—
Moz1 has entered CATL’s power-battery PACK production line, where it performs non-standard processes such as inserting high-voltage testing plugs. Its task success rate remains stable at above 99%, and its daily workload can reach three times that of a skilled worker. Qianxun Intelligence is also conducting scenario-based collaborations with companies including JD.com, Bosch, and Schaeffler.
Another model, Moz2, is attempting to transfer capabilities from the same-structure base model to public-service scenarios such as supermarkets, hotels, and office buildings. At WRC, it also demonstrated natural voice interaction for the first time, supporting human-like interactions such as waving, making a finger-heart gesture, and shaking hands.

Looking back, it becomes clear why Qianxun Intelligence chose to run the same demo again at two different conferences.
In the previous stage of embodied intelligence, the question most easily noticed was whether robots “could do it.” But as more and more robots begin performing these actions, the questions for the next stage have changed.
Can the robot perform the same task continuously? Can it still do so when the environment changes? After a failure midway, does the system know what went wrong? Can the problem encountered this time enter the next training cycle, so that the next version makes one fewer mistake?
Ultimately, the end goal of embodied intelligence has never been to build an all-powerful robot. It is to establish a technical system capable of growing on its own.
The “time density” mentioned above is therefore not simply measuring how many percentage points a task’s success rate has improved over a given period.
It is actually evaluating how quickly a technical system can process real-world failures: whoever can absorb problems faster and return improvements to the robot faster will expand the boundaries of its capabilities more rapidly.
This is becoming an important benchmark for assessing the long-term potential of embodied-intelligence companies.