A Robot Takes a Go-Kart Through Consecutive Turns—Fully Autonomous, All in One Take
While scrolling through social media at home over the weekend, I came across a video that absolutely amazed me.
A robot driving a go-kart was racing around a track—taking corners, accelerating, and steering, all in one continuous shot.
I never expected our robots to be showing off on a go-kart track already!

Wait.
This doesn’t look anything like what I’ve seen in the WRC!
Robots at exhibitions run around with sparks and lightning effects, but I never imagined that the robot in this video had secretly enrolled in racing school behind humanity’s back.
It was a complete veteran driver: high-speed cornering, agile obstacle avoidance, and a flawless lap from start to finish.
Compared with previous demos, where robots slowly walked over or bent down to pick things up, the robot in this video actually sat down in the vehicle on its own, synchronizing its hands, eyes, and feet, while performing multi-contact-point balancing and precise force control at full speed.
I’ve posted the video here—take a look.
This kind of bipedal humanoid robot can coordinate its eyes, hands, and feet across its entire body, carrying out detailed operations during high-speed movement and complex changes in posture. Such a performance has hardly ever appeared in previous humanoid-robot demonstrations.
The company that released the video is called Gongsheng Zhixing.
Gongsheng Zhixing is developing an end-to-end integrated perception-and-control “brain” for bipedal humanoid robots—in other words, a foundation model.
Founder Ding Pengxiang was born in 1996. Every other member of the core team was born in the 2000s, and all are currently pursuing PhDs. The company was founded just two months ago.
As for this seemingly somewhat “misguided” demo, Ding Pengxiang explained:
I hope robots can be more like humans, rather than only performing mechanical service functions. Humans can drive go-karts, so robots should be able to do it too.
That raises an obvious question: Why would a group of young people who have not even received their diplomas dare to take on the most difficult and cutting-edge track in embodied intelligence?
Betting on the Hardest Path: They Studied Hierarchical Systems—Now They’re Taking on End-to-End Models
How new is the field of “end-to-end” bipedal humanoid-robot models?
New enough that there are no ready-made candidates on the market.
According to Ding Pengxiang, this direction has only become popular over the past two years. People who understand the technology are still pursuing their PhDs, so the team has had almost no way to find qualified candidates through conventional talent pools.
Those who graduated years ago cannot keep up with the pace, while those without PhDs often lack a sufficient understanding of the underlying technology. The most cutting-edge breakthroughs are currently in the hands of this group of top PhD students.
The team at Gongsheng Zhixing is made up of precisely this kind of young, cutting-edge talent.
They have won two of the most prestigious academic honors in China’s embodied-intelligence field—two Best Paper awards—and had another paper selected as a Best Paper Candidate. They also built China’s first embodied foundation model to surpass 2K GitHub Stars.
More importantly, they were not followers who joined after the trend took off. They were among the earliest Chinese teams to explore embodied foundation models, and are also some of the most committed long-term players in the field. The team has published more than 40 papers at leading conferences, with technical coverage spanning the full stack of perception, decision-making, control, data, and systems. A series of industry-first projects—from dual-system architectures and lightweight models to multi-configuration systems—have all come from this young team.

Their view is clear: the form in which robots will ultimately enter human life most extensively is the biped.
Following the evolutionary paths of autonomous driving and large models, end-to-end is the most efficient technical paradigm.
This judgment comes from their firsthand research experience.
As early as December 2023, Ding Pengxiang began equipping quadruped robots with a “brain.”
As first author, he published QUAR-VLA, the first VLA task paradigm designed for quadruped robots. It assigns the upper-level model—the brain—to understand vision, language, and task intent, and then passes decisions to the lower-level motion system—the cerebellum—for execution.
This approach of having “the brain handle decisions and the cerebellum handle motion execution” is what the industry calls a hierarchical architecture. Companies such as Figure are now following the same approach.
In subsequent research, however, he gradually began to identify its problems.
He describes hierarchical systems as “theoretically feasible, but not particularly elegant”:
The brain and cerebellum each do their own thing. Whenever a new task is introduced, the upstream output must be fine-tuned specifically for that task before the system can achieve good results.
The deeper issue is an inherent structural flaw.
The brain and cerebellum are trained separately and only connected during deployment. Even if each component performs better individually, there is no guarantee that the overall system will be optimal.
Every translation between the two layers introduces another cascade of errors and another loss of information.
Ding Pengxiang concluded: “As long as the system is hierarchical, there will be an information bottleneck in the middle. This interface design means that a hierarchical architecture cannot achieve Scaling in the true sense of the word.”
This conclusion also came from his long-term observation of the development of autonomous driving and large models.
Early autonomous-driving systems relied on the coordination of multiple modules for perception, prediction, planning, and control. Later, researchers began exploring ways to let models learn driving capabilities directly from data.

After Tesla’s FSD shifted to an end-to-end approach, its performance improved dramatically.
In Ding Pengxiang’s view, robots may undergo a similar transformation, with models learning the complete “perception–understanding–action” process from data.
His reasoning is that, since 2023, the hierarchical approach has been repeatedly validated. If the future scaling of data will ultimately push systems toward end-to-end models, there is no need to take a detour first.
That shift is taking place in 2026.
Google’s Gemini Robotics 2 has already unified full-body movement—from the feet to the fingertips—within a single policy.
Yet the consensus at RSS 2026 states that a single end-to-end model cannot cover complex full-body dynamics, and that hierarchical modularization is currently the optimal deployment strategy.
Gongsheng Zhixing has bet on pure end-to-end, a more aggressive path followed by far fewer people.

Why bipedal robots specifically? Ding Pengxiang’s answer is first principles: the buildings, tools, and environments of human society are all designed around the human body’s physical configuration.
Wheeled robots can only work in fixed settings such as production lines and shopping malls. Bipedal robots can go outdoors, climb stairs, drive cars, and perform different tasks across environments.
Generality means costs can be spread across more applications, while bipedal robots also offer the approachable familiarity of a human-like form.
Timing is equally important. Before April 2026, bipedal humanoid robots did not even have a general-purpose teleoperation model. Without a source of data, it was impossible to train a foundation model.
It was not until NVIDIA open-sourced Sonic that the industry had a foundation for producing data at scale.
The field is no longer in its early stage of “wanting to train a model but having no data,” yet it has not matured to the point where the competing technical paths have fully converged. Gongsheng Zhixing is betting on this window in between.
The Technical Divide: While Others Learn Kinematics, They’re Tackling Dynamics
First, it is important to understand a key divide: most robot foundation models today are actually learning kinematics.
Kinematics focuses on moving from point A to point B, determining where a hand should move, and whether an object has been grasped. This works fine for fixed-base robotic arms because they do not fall over.
When a bipedal robot reaches for an object, its center of mass shifts at the same time.
When it squats to pick something up, its waist must lean forward, while its ankles and legs redistribute force to maintain balance.
Once friction, collisions, inertia, and contact are added to the equation, the model is dealing with a complete full-body dynamics system.
According to Ding Pengxiang, a G1 bipedal humanoid robot has 29 degrees of freedom, making it far more complex than a robotic arm with seven degrees of freedom.
In the past, robots learned “motion trajectories.” Humanoid robots must learn “how a body acts in the physical world.”

The shift from kinematics to dynamics—from completing tasks to maintaining stability—is the true technical divide for bipedal-robot models.
This explains why Gongsheng Zhixing has taken its model all the way down to Joint Target, or the joint-target layer.
In a hierarchical architecture, the brain first outputs kinematic targets, and the downstream cerebellum then converts them into joint actions.
Gongsheng Zhixing removes this intermediate step, allowing the model to directly handle the body states of dozens of joints.
That raises another question: If the cerebellum is removed, can the robot still remain standing?

With imitation learning alone, the model only learns standard movements. It has not encountered the various unsteady body states that can occur during execution, so it may simply fall when faced with an unfamiliar state.
What it lacks is the ability to recover and stand steadily after failure.
At the core of Gongsheng Zhixing’s approach is a dual-domain collaborative optimization mechanism based on “task-behavior modeling and motion-prior distillation.”
One optimization pathway retains behavioral cloning to ensure task accuracy and motion-fitting capability. The other uses DriftDistill to convert the stability, disturbance resistance, and Failure Recovery capabilities accumulated by low-level controllers into an intrinsic motion prior for the unified model.
This is not a simple matter of adding two Losses together. Instead, it integrates two types of capabilities—“completing tasks accurately” and “controlling the body stably”—into a single model, allowing the large model to acquire both task intelligence and physical intelligence at the same time.
If DriftDistill can continue scaling with the model and the amount of data, the larger problem it is trying to solve is this:
Beyond cognitive Scaling, can robots Scale their motion-control capabilities as well?
Ding Pengxiang calls this motion-capability aggregation module the Motion Expert, which is currently close to 1B in scale and continuously absorbs capabilities from different motion-control models through distillation.

Going one step further, end-to-end large models must also solve the problem of how much force a robot should apply.
Today, many robotic systems primarily control position. But once robots truly enter the real world, being in the correct position does not necessarily mean the task is complete.
A hand may reach a drawer handle without knowing how much force to apply, leaving the drawer unopened. When handing something to a person, the position may be correct, but excessive force can still be unsafe.
Gongsheng Zhixing’s technical route is to first use Force Expert as a safety and contact expert. It learns force application, compliance, impedance, and contact feedback, and is then gradually integrated into Motion Expert and the final end-to-end model through policy distillation.
Position determines whether a robot can “get there”; force determines whether it can truly “get the job done.”
This is also what Gongsheng Zhixing means by a “full-body physical-world model”: enabling the model to understand how the body experiences force, loses balance, and makes contact, and then directly turning those constraints into actions.
Data, Emergence, and a Cool-Headed Timeline
When it comes to actually training models, the first obstacle is the data shortage.
As of early 2026, the world had approximately 500,000 hours of compliant robot data—less than one twenty-thousandth of the data available for large language models.
The awkward reality is that much of this data captures stationary operations: robots working in place inside confined workspaces. It is of limited help for full-body mobile manipulation.
Gongsheng Zhixing’s blog has released more than 10,000 hours of data, including several thousand hours of self-collected data.
What full-body coordination models truly lack is strongly coupled full-body mobile-manipulation data from real life, such as riding a bicycle, driving a go-kart, or inflating a basketball.
Their solution is TrajBooster: extract end-effector trajectories from robotic arms or wheeled robots, then have humanoid robots track those trajectories and maintain balance in simulated environments. This brings actions from different sources into a shared space, reuses the task diversity already defined for robotic arms, and provides low-cost data for pretraining.
When discussing model capabilities, Ding Pengxiang rarely uses the word “generalization.”
As he puts it: “The word has very little substance. It is as empty as saying someone is a good person.”
What he wants is “emergence”: at the joint-angle level, for example, if a model is trained on walking and jumping data, it should be able to combine them into an entirely new skill—running—during testing.
That is because only at the joint level can certain motion capabilities potentially be combined into genuinely new actions.

His timeline is notably clear-eyed.
He estimates that around 100,000 hours of data will still be used primarily for laboratory-level demos. Truly meaningful large-scale commercial applications may not arrive until after the million-hour mark.
The industry has not even converged on the most suitable form of data. Approaches involving electromyography, exoskeletons, motion capture, and Pico teleoperation are still competing.
In his view, the three hardest problems currently facing humanoid robots are, first, how perception and control should ultimately be combined; second, what type of data should be used; and only third, the scale of the data itself.
Ding Pengxiang hopes that one day, a robot-model architecture designed by a Chinese team will also be widely adopted by peers in North America.
We do not want to be Followers. We want to lead the development of the technology.
Now, back to that go-kart.
The people sending it through high-speed turns are a group of young people who have not yet graduated.
They are betting on the most difficult and cutting-edge direction in embodied intelligence.
Their wager is that end-to-end will eventually replace hierarchical systems.
The industry is still far from proving that the final path has already been established. Ding Pengxiang himself even acknowledges that the entire bipedal foundation-model industry has yet to truly converge.
But the cutting-edge nature of a startup team is sometimes defined not by how quickly it can implement a mature answer, but by whether it can spot the next truly important question earlier than its peers—and have the courage to start solving it while the answer is still unclear.