Today’s embodied AI companies often film their demos to look polished and exaggerated, almost like advertisements. But have you ever seen a genuine, unedited video?

A few days ago, someone in the industry sent me a 10-minute demo: one continuous shot from start to finish, with no editing, no off-camera remote control, and no human instructions. It even looked a little rough around the edges. But what the video showed left me slumped in my chair, unable to calm down for a long time (doge).
In the video, a robot extends its arm out of a window to clean the glass; when it cannot reach a higher spot, it knows to bring over a box and stand on it; and when a long-horizon task is interrupted midway, it can resume it precisely where it left off…
Even more astonishingly, the video features robot hardware from both Unitree and AgiBot. These two “competing” platforms differ completely in hardware architecture, degrees of freedom, and sensing systems, yet they share a single “brain,” cooperating and working closely together. This is a rare example anywhere in the world of a general-purpose brain operating across embodiments.
Without exaggeration, this may be a demo capable of reshaping the global embodied AI industry’s understanding of the field and overturning prevailing judgments about its technological trajectory. It might even change the Scaling Law…
So where did this “miracle team” come from?
A Frame-by-Frame Breakdown of the Video—One Jaw-Dropping Moment After Another
There is simply too much information packed into this video. After carefully analyzing it frame by frame, I found myself becoming more and more fascinated.
- Working in a confined space, with astonishing stability
The video was filmed in a rental apartment of roughly 15 square meters. It is packed with furniture, the passageways are narrow, and even turning around is difficult. In such an environment, teleoperation is basically impossible—there is not enough room for an operator on site—and it would also be extremely difficult to script the entire sequence in advance.

The environment design alone seems to be telling us directly: this can only be a video of robots learning, evolving, and making decisions entirely on their own.
The two robots perform household chores in parallel within the same space, sensing the boundaries of their surroundings in real time and planning their paths.
There are no collisions, no getting lost, and no pauses throughout the entire sequence. They adapt perfectly to an unfamiliar and complex real-world environment. Even the “battle-damaged” robot tethered by a leash demonstrates remarkable agility.
There are no cuts, reshoots, or human instructions anywhere in the video.
- Reaching out of the window to clean the glass—the ultimate embodiment of dynamic control and autonomous reasoning
From the very first task, the robot displayed a subtlety that went far beyond what I had imagined.
A Unitree robot uses a squeegee to clean the glass. Anyone who has used one knows that too little force will not clean the glass properly, while too much force produces an unpleasant noise. Compared with wiping glass with a cloth, this requires much more precise force control and spatial awareness. Right from the start, the robot voluntarily made the task more difficult for itself.
After a few strokes, it noticed that one spot had not been cleaned properly and independently concluded that the dirt might be on the outside of the window.
It then performed a sequence of movements I had never seen before: turning sideways, leaning back, peeking out, and extending its arm until the hand holding the squeegee reached outside the window. (Steady as a rock.)
For some existing embodied models, achieving reliable force control is already the limit. But this robot seamlessly combines precise torque control with awareness of the surrounding space—it did not hit the window frame while extending its arm—as well as dynamic computation.
It has not merely reached an industry-leading level in each individual capability; it has demonstrated their seamless integration.
What left me most stunned was that, after realizing it could not clean the glass properly from inside, it immediately reasoned through the situation and decided to extend its arm outside the window. This was the first time I had seen a model reason and make an autonomous decision in a way that resembled a human.
This shows that the model can fuse visual, tactile, and dynamic information into a unified perception system, while also possessing the ability to evolve on its own. For the first time, I truly believed that robots can really do useful work—and that in the future, they may do it better than humans.
- Closing the loop on long-horizon tasks, without error accumulation
After cleaning the glass, the Unitree robot smoothly returned the squeegee to its original position. Its placement was precise and the movement seamless.
Do not underestimate this finishing move. It means the robot completed the entire loop of “retrieve tool—perform task—put tool away.”
Nearby, the AgiBot robot approached the washing machine, took out a clean cushion, and placed it behind the sofa. It then returned, took out a clean stuffed toy, and placed it precisely inside the second level of the wardrobe. Every item went back where it belonged.
The AgiBot robot accidentally bumped into the glass, which actually makes it even clearer that the process was being controlled autonomously. After all, cameras cannot always resolve the problem of reflections on glass; with human teleoperation, the obstacle would usually have been avoided in advance (doge).

These two scenes represent only the most basic capabilities displayed in the 10-minute video. Across the long-horizon tasks, every action is smooth, precise, and completed correctly on the first try. The robots do not accumulate more and more errors as the tasks become longer and more numerous.
This addresses the long-sequence task problem that traditional VLA models fear most.
Judging from the movements alone, this model may well employ an entirely new architecture. Across an exceptionally long 10-minute task chain, it continuously corrects itself and produces stable outputs, keeping every action precise and controllable. Its robustness appears to be among the best in the industry.
- Tasks can be interrupted at any time, and the model can resume from the interruption point
Next, the sound of an alarm clock appears in the video. (Listen carefully with the volume turned up—the sound is not particularly obvious.)
At first, I did not understand the purpose of the alarm. Only after watching it repeatedly did I realize that it seemed to be some kind of preset reminder, interrupting the two robots’ current tasks and prompting them to perform a special task: tidying the desk and organizing the refrigerator.
The remarkable part is that after finishing those tasks, the robots resumed the work they had been doing before the interruption. This indicates that the model can handle arbitrary interruptions, resume from a checkpoint, and recover pending tasks.
Compared with demos on the market that can complete only one task at a time and must be handled with extreme care, this seemingly rough video shows something very different: in a real working environment, the robots can not only keep working continuously, but even “multitask.”
Next comes an entire household organization routine: the Unitree robot picks up an organizer bag and places it on the table. After noticing food on top of the refrigerator, the AgiBot robot determines that it should be refrigerated, puts it away, and then takes out food that needs to be thawed. (I suspect this mysterious team may be hinting that robots will soon be able to prepare a full four-dish-and-one-soup meal.)
There are no step-by-step instructions. The robots decompose the tasks, plan their movements, and complete the entire process from retrieving items to putting them away.
Today, most robots can execute only a single task in a serial workflow. If anything disrupts them midway, the task usually collapses. This model, however, appears to possess something resembling human judgment of task priorities and dynamic scheduling capabilities, allowing it to switch between task flows at any time. The intelligence demonstrated by this flexibility and stability far exceeds that of most known models.
- AgiBot puts a scarf on Unitree—a defining moment in cross-embodiment collaboration
Then came the climax of the entire video.
The two robots seemed to reach an agreement: they decided to organize all the clutter on the table in one go, avoiding the need to carry things back and forth.
The Unitree robot kept placing objects on its body until both hands were occupied and it no longer knew how to proceed. At that moment, the AgiBot robot gently walked over, picked up a scarf from the table, and slowly draped it around the Unitree robot’s neck.
The AgiBot robot noticed that the scarf was long, so it lifted it with both hands. Unitree seemed to understand what it intended to do and proactively bent down. AgiBot then brought the scarf over Unitree’s head and placed it around its neck.
Come on, seriously?! The movement was so fluid, and the eye contact between them so expressive—how is there even a hint of a couple’s chemistry here?
This was not a preprogrammed division of labor. It was a collaborative solution autonomously discovered by the model across two different embodiments. The two robots, made by different brands, were able to independently assess the limits of their respective capabilities and complement each other perfectly. This is already remarkably close to human collaboration—and without any verbal communication, they even seemed more in sync.
This may be the world’s first instance of cross-embodiment interaction. It may also be the first model capable of generalizing across different embodiments.
My guess is that this mysterious team is using two robot platforms that are effectively competitors to tell everyone: this is what a truly general-purpose brain looks like.
- Stacking towels and carrying slippers—can robots not only work, but also “take shortcuts”?
After the Unitree robot slowly walked away with objects hanging all over its body, the AgiBot robot continued organizing the remaining items.
At one point, it picked up a towel and tried to hang it on itself. After one attempt, then a second, then a third, it finally managed to drape the towel over its shoulder.
Three shocks in a row!
The robot seemed to know what required the least effort: draping the towel over its shoulder was easier than holding it in its hand.
The model can learn autonomously, continuously improving after one or two failed attempts until it succeeds.
Its understanding of its own embodiment may be precisely what enables it to operate across different embodiments.
And this was not an isolated case.
Look closely: when organizing the slippers, the robot carries them by the hanging strap of the new slippers rather than by the shoes themselves, because carrying them by the strap requires less effort.
Shocking again! While other robots are still carefully designing movements just to get work done, this model not only completed the task—it even learned how to “take shortcuts.” That is remarkably intelligent.
- Another climax: the robot learns to use a tool and tries to stand on a box to reach higher
But the story was not over. The next scene shocked me the most.
While putting away the objects hanging from its body, the 1.3-meter-tall Unitree G1 robot was stumped by the problem of placing the scarf on the third shelf of the cabinet.
Then something remarkable happened. It did not freeze or give up. Instead, it reasoned like a human and found a box on its own. That’s right—it found a box and tried to stand on it to gain enough height to finish putting the scarf away.
It located the box nearby and pushed it onto the floor. It tried to bend down and pick it up, only to discover that it could not bend over far enough.
After several attempts, just when I thought it would finally give up, it kicked the box toward the cabinet.
The entire process demonstrated:
Self-directed reasoning and decision-making. The robot knew how to reach an object positioned high up, and it also knew how to move the box when it could not bend down to pick it up.
Continuous exploration of its physical capabilities. As it repeatedly tried to bend down, the robot seemed to gain an increasingly precise understanding of its physical limits and made choices suited to its capabilities.
A process of self-evolution. Without any instruction, the robot learned to kick the box with its foot. It was able to make use of its body and interact with the environment, continuously adjusting its strategy along the way.
That scene genuinely sent a chill down my spine.
After all, the autonomous use of tools is one of the defining capabilities of human beings.
And the robot in this video really does seem to have learned how to use tools: it knows that the box can bear weight, that standing on it can increase its reach, and even that when it cannot bend down, it can kick the box into position with its foot.
Behind this sequence is a deep understanding of the rules of the physical world, continuous exploration of its own capabilities, and the combination of autonomous decision-making with evolving strategies.
- Complementary capabilities across embodiments and collaborative work—from individual intelligence to collective intelligence
Unitree kicked the box over, but it landed crooked. Just as it was considering how to straighten it, the 1.7-meter-tall AgiBot Expedition A3 walked over.
It seemed to recognize Unitree’s determination and predicament, put down the small cart it was carrying, and proactively switched tasks to help its companion.
Unitree bent forward and lowered its head as if receiving a graduation medal, while AgiBot slowly removed the scarf and placed it on the shelf.
Every detail of this sequence demonstrates capabilities that distinguish the model from others.
The two robots’ understanding of and coordination with their own hardware capabilities—as well as each other’s—has gone beyond individual intelligence and moved toward collective intelligence.
AgiBot brought the scarf around the back of Unitree’s head and deftly removed it, demonstrating an extraordinary understanding of force control and spatial awareness. (At this point, you can imagine how other robots might have yanked the scarf off.)
The robot folded the scarf into three sections. This was an autonomous judgment about how best to fit the scarf into the wardrobe. This is what genuine embodied “intelligence” looks like.
After helping Unitree, AgiBot quietly walked away and placed the last piece of clothing securely into the washing machine. (I suspect it concluded that the item draped over the washing machine was dirty laundry.) And with that, the 10-minute, one-take video finally came to an end.
Although I do not know the specific architecture of this model, I believe everyone watching was left as stunned as I was.
The first instance of cross-embodiment collaboration; the first time seeing a robot evolve autonomously; the first time seeing a robot make human-like autonomous decisions and choose the optimal path; extraordinary spatial awareness and force control; tasks that can be interrupted at any time; and the ability to learn how to use tools…
Every action and every frame in this model could have served as an excellent demo on its own. Yet all of them appeared quietly in this simple, direct, 10-minute single-take video.
It was as if everything came effortlessly and naturally, conveying the calm confidence of a champion.
The Underlying Technology Breaks Free from the Constraints of Mainstream Models
Why were the robots in the video able to deliver so many unforgettable moments?
I learned from an informed source that the most important reason is that this model has broken free from the established framework of today’s mainstream embodied models.
At this stage, mainstream embodied models can roughly be divided into three categories: VLA, WAM, and traditional world models. All of them face bottlenecks that are difficult to overcome for now.

△ Image generated by AI
More specifically, VLA is the “data-and-intuition school.” It relies on end-to-end fitting over massive amounts of visual, language, and action data. But at its core, it is “looking at images and guessing the action,” without genuine physical reasoning. In long-sequence tasks, errors continue to accumulate; its ability to generalize to unfamiliar scenarios is almost nonexistent; and it is highly tied to a specific embodiment.
WAM and traditional world models belong to the “speculative prediction school.” They attempt to model the laws of the world first and then plan actions, but suffer from the fatal problem of a disconnect between knowledge and action. The model may understand a scene and predict its state, but when faced with real operations requiring physical and dynamic reasoning, it still has to rely on training data to simulate what might happen and cannot adapt to dynamic changes in the real environment.
These mainstream embodied models share a common fundamental flaw: data-driven task fitting.
In other words, all actions are based on “probabilistic guesses” derived from massive datasets, rather than on a deep understanding of physical rules. As a result, the upper limit of the model’s capabilities is firmly constrained by its training data.
I heard that the model in the video was trained using only a few dozen hours of video data. If such a small amount of data can produce results this astonishing, then does Scaling Law still exist? That question mark is certainly justified.

Judging from the details, the model in the video demonstrates at least four capabilities that current VLA and world models have yet to match:
-
Dynamic modeling: Its movement trajectories satisfy dynamic feasibility constraints. It can calculate compensating torques and feedforward terms, giving it impressive extrapolation and generalization capabilities;
-
Self-awareness: Unlike VLA, which relies on conditioned reflexes, and WAM, which relies on temporal statistics, it can think like a human about “which action will bring me closest to the goal next”;
-
Adaptability: It can fill in gaps in causal reasoning, address long-tail problems, and continuously adjust and evolve while carrying out tasks;
-
Self-evolution: Its skills improve as it completes tasks, its strategies continue to evolve, and it seems capable of generating its own learning objectives and actively reasoning toward them.
Three foundational technologies support all of this.
Physics-constrained dynamic learning. Unlike the purely data-driven training paradigms of VLA and WAM, this approach places physical rules at the forefront and uses dynamic prediction to drive behavior. Carrying slippers by their straps, draping a towel over the shoulder, and standing on a box to reach higher are not behaviors directly taught by data. They are optimal solutions the model discovered through interaction with the environment.
Unified cross-embodiment modeling. Through a unified action representation space and embodiment-adaptation mechanisms, the same model can operate on hardware platforms from different brands with different architectures. This effectively decouples the software and hardware of embodied intelligence, bringing an end to the era in which algorithms were bound to specific hardware.
Robust long-horizon closed-loop control. With a global task-scheduling framework, the model can autonomously handle unexpected disruptions, action errors, and task switching throughout the process. It avoids the collapse caused by accumulated errors in traditional models during long workflows, providing the core capabilities needed for deployment in real homes and complex environments.

If you are familiar with the current state of embodied models, you were probably left speechless after reading this.
While other teams’ demos are still stacking up short runtimes, heavy post-production, and single-task scenarios, this mysterious video has already demonstrated a real-world deployment featuring 10 uninterrupted minutes, multiple brands, multi-robot collaboration, end-to-end execution, and self-evolution.
While other models are still “guessing actions from data,” this model seems to have achieved genuine self-directed decision-making, self-learning, and self-evolution. It shows us that robots can not only really do useful work, but that their capabilities are rapidly evolving on their own, leaving ample room for the imagination.
More importantly, this model seems to show us the future everyone has been discussing so intensely—the day, supposedly three, five, or even ten years away, when robots can truly replace humans in carrying out work—already unfolding before our eyes today.

An as-yet-unknown team seems to have overturned every established technological path, overturned Scaling Law, and created genuine embodied “intelligence,” bringing the field close to its ChatGPT moment.
We still do not know who created this mysterious model, but I am sure everyone is asking the same question I am right now:
Who on earth is behind it?