Skip to content
Main Site News Console

Mysterious Embodied AI Team Unveils Major Demos, Revealing Its Self-Evolving Model Technology Roadmap

· 量子位
国内AI

Do you remember that mysterious, 10-minute-long continuous-shot video that was making the rounds in the embodied intelligence community last week?

As soon as the video came out, netizens were saying: Unbelievable—it looks like something from outer space…

Some even gave it the ultimate compliment: This has to be AI-generated, right?!

People in the industry were scrambling to find out which company had made it. Word has it that even a major company held a special meeting to study it… Our own inboxes were flooded with questions from every direction.

After watching the 10-minute continuous-shot video, everyone began saying that the ChatGPT moment for embodied intelligence might really have arrived. Some called it an important milestone for embodied intelligence; others described it as an architectural innovation. That’s because it showed that:

  • Robots possess a certain level of cognitive ability, rather than merely making probabilistic predictions;
  • Robots can understand the logic behind human behavior, rather than simply overfitting;
  • Robots are capable of continuous self-evolution, producing many astonishing moments.

Of course, because the robot’s performance was so “sci-fi,” some people voiced their doubts: Could this be teleoperation? (Although the team said they genuinely had no idea how teleoperation could be carried out in such a small space.)

The mysterious team had a rather unusual way of responding to these doubts: They released several more demos—still rough, still continuous-shot—as if they were trying to say:

“Intelligence is coming, whether you believe it or not.”

Because of our previous article, we were able to get in touch with this mysterious model team and ask the obvious question: Why not simply come forward and claim credit?

Their response was: The model is like a person—let it speak for itself.

We learned that the mysterious model’s internal codename is “MVP,” meaning that the model’s capabilities are like those of a real person (Make Veritable People). The team also said it could be translated directly as “Be a Person” (doge).

The “Be a Person” large model.

Judging from the performance in these newly released demos, the robot has indeed “become a person”: It thinks like a human, makes decisions autonomously, moves fluidly, and behaves logically. According to the team, they are about to let the robot take to the streets and demonstrate its capabilities directly to the public in an open environment.

Video 1: Unified hand-eye coordination and an impressive grasp of dynamics

http://www.qbitai.com/wp-content/uploads/2026/09/飞书20260902-135110.mp4

If you were holding a bottle of water and someone poked your arm, could you react instantly without spilling a drop?

The robot holds a cup full of water in its left hand. When a human (rude) stick jostles it, the robot shifts to the other side, then smoothly returns to its original position without spilling a single drop.

If you were holding a bottle of water and someone poked your arm while also knocking over a soda can beside you, could you react instantly—keeping the water from spilling while making sure the can did not fall?

A human then pushes three soda cans on the table from the right, and the robot’s right hand somehow catches the cans at the same time.

The robot performs like a tai chi master: It knows how to “neutralize force,” dissolving the opponent’s impact almost imperceptibly, while also understanding how to apply just the right amount of force.

And if you watch closely, you’ll notice that when the cans are pushed, the robot’s left arm has not yet finished stabilizing the cup of water, yet it still uses its right hand to steady the cans.

That moment is impressive not only for a robot—it has already surpassed many people. Try it yourself and see whether you can perform both tasks simultaneously with two hands. (doge)

Mainstream VLA and WAM models have a problem: They tend to fail when faced with task scenarios involving occlusion, contact, deformation, and other situations that require physical reasoning. That is because they merely imitate the outward form of an action without understanding the physical principles behind it.

The “Be a Person” model, however, appears to have equipped the robot with a “physical brain.” Only by unifying dynamics prediction with long-term state representation can trajectories naturally satisfy dynamic feasibility.

Judging from the robot’s reaction speed and fluidity of movement, its actions no longer seem to be “guessed” through rote memorization of data. Instead, they appear to be “learned” like a human’s, through an understanding of physical rules.

The difference is roughly that between “copying a drawing by following its outline” and “understanding why it works.”

Video 2: Understanding space and object properties, and learning human habits

In the second video, the humanoid robot officially takes over the housework. The “Be a Person” model team also says that this video shows the robot completing the task in a zero-shot manner. (If true, that level of generalization is incredibly strong!)

They say the model’s zero-shot success rate has already reached 80%.

I’ve named this video “The Obsessive-Compulsive Robot: The Living Room Must Not Be Messy” (doge).

The robot first returns the small stuffed toy in its hand to the storage stand by the entryway, following the principle of “putting things back where they came from.” It then drags a storage basket filled with miscellaneous items over to itself, bends slightly, and takes out a black hat and a purple hula hoop one after another, hanging them neatly on the coat rack.

Finally, it removes a large flower-shaped cushion from the basket. Clearly, it is capable of bending over, but this time it simply tosses the cushion aside, and it lands on the sofa.

Throughout the process, the robot’s movements are not particularly nimble, but it gives the distinct impression of being completely in control.

The robot’s understanding of space and object properties is remarkably human: It knows that objects can be dragged across a smooth floor, that ring-shaped items such as hats and hula hoops can be hung up, and that a flat cushion will unfold on its own.

It even knows how to “take shortcuts”: If it can throw something directly, it will not go to the trouble of bending over. (Lazy people really do drive scientific progress, while lazy robots drive the evolution of intelligence.)

Mainstream models have yet to achieve human-level generalization. VLA relies on brute-force computation in the hope that “enough force can work miracles,” and tends to fall apart in unfamiliar environments. WAM, meanwhile, pieces actions together like building blocks; it learns too rigidly and stops working as soon as the situation requires even a little flexibility.

The robot running the “Be a Person” model, by contrast, appears remarkably stable. It combines an understanding of space and object properties with an understanding of human habits, successfully completing a long-horizon task.

Video 3: Stretching, shaking, and laying out a bedsheet—the robots are working as a team

When we were children, our elders would often ask us to shake out freshly washed bedsheets.

This video features two different models of robots. Reportedly, all their training was completed by watching human videos, so they once again demonstrate cross-embodiment collaboration—the same brain in different bodies—and genuinely work together as a team.

A bedsheet that appears to have just come out of a dryer is held at two corners on one end by a taller robot, while the two corners on the other end are held by a shorter robot.

The shorter robot takes a step forward and bends down, pulling the bedsheet taut. The taller robot then gently shakes the sheet several times with both hands.

Although the robots’ postures in the video are still somewhat stiff, their behavioral logic is remarkably human! How is this different from shaking freshly washed clothes a few times to prevent wrinkles?

It is worth noting that mainstream models have a fatal flaw: Unlike humans, they cannot truly self-evolve. VLA can only make minor patches and adjustments without changing its “neural circuitry”; WAM is too rigid in how it performs tasks and cannot adapt when the environment changes.

The MVP model is different. It appears to unify an understanding of the physical world with an understanding of the logic behind human behavior.

Video 4: Energy-driven behavior—the optimal-energy solution

When a robot learns to think and realizes that carrying everything away in one trip is less effort than organizing things over multiple trips, what will it do?

In the video, the robot organizes objects much like a human would: No matter how many items there are, it tries to avoid making a second trip.

It first pulls a black scarf from the basket and drapes it around its neck like a shop assistant. Next, it picks up a purple ring, loops it around its forearm, and even raises its hand so that the ring slides down to its elbow, successfully transforming itself into a walking coat rack.

Then come the slippers and earphones. After all this, the robot is covered in items, yet it calmly turns and walks away, leaving behind a graceful silhouette.

In the one-minute video, the robot puts away four items. It not only demonstrates the ability to handle long-horizon tasks, but also shows an understanding of each object and the ability to establish a mapping between each item and its own capabilities.

Today, mainstream models performing the same task would most likely carry the objects one by one, because they are still essentially making guesses based on statistical fitting. They lack a human-like understanding of the world, as well as autonomous intent.

If the robot in the video saw this, its inner monologue might be: Seriously, bro? You’re already making the trip—how can you leave with so little? Wouldn’t it be easier to take more at once?

When making decisions, the robot running the “Be a Person” model seems not to need to deliberately “think” about what to do next. Instead, it simply glides naturally in the direction of decreasing energy, like water flowing downhill.

In this process of “going downhill,” thinking, exploration, and action happen as one seamless whole.

Perhaps this is what true world understanding and autonomous intent look like.

Summary

After watching these four demos, I finally understand why the model is called “Be a Person.” Every movement the robot makes really does resemble a human decision-making process: In long-horizon tasks, noisy environments, and collaborative interactions, it demonstrates continuous learning and evolution in its understanding of the physical world, the logic of human behavior, and its own capabilities.

The robot can understand space and object properties like a human. Its thought processes and operating habits when doing housework resemble those of people, and it can also self-motivate, learn autonomously, and continuously self-evolve like a human.

What’s more, the robot displays the following traits:

  • Consistency of movement: The robot is like a “mature human”—logically coherent, fluid in motion, physically knowledgeable, and equipped with strong hand-eye coordination.
  • Continuity of thought: The robot does not “give up halfway,” and it becomes “more reliable with use.” Its long-sequence tasks unfold in tightly connected stages.
  • Even a distinct personality: Robots running the “Be a Person” large model seem to have their own unique “style of doing things,” as though they already possess individual personalities.

There are many other characteristics as well. But I am genuinely beginning to believe that the moment when embodied intelligence can truly serve humanity is drawing closer and closer.

Intelligence may really have emerged—it may simply be happening in places we have not yet seen.