Robots Can Learn by Watching a 3-Second Demonstration
Mengyao, reporting from Aofeisi
I’m telling you, this is wild… The GPT-3 moment for robotics may have really arrived!?
GPT-3’s ability to learn a new task from just a few examples is now playing out on robots—
Even for a task they’ve never learned before, all you have to do is show a robot a 3–12-second demonstration, and it can immediately perform the task itself!

It can even generalize from what it has learned.
First, you teach the robot to sweep things with a brush. Then you replace the brush with a banana—and it actually picks up the banana and keeps sweeping as if it were the most natural thing in the world—

Still not impressed? Then let’s give it a dustpan too~
Look at that. This time, nobody even demonstrated the complete movement, yet it figured out how to coordinate both hands on its own: one hand sweeps, while the other catches the debris. Incredible:

This is GEN-1.5, the newly released robot foundation model from Generalist AI.
It’s a robotics model focused on One-shot Learning. After pretraining on large-scale physical data, it enables robots to learn new tasks quickly by watching just a few seconds of demonstration.
What’s more, the model can combine two different teaching demonstrations to learn a task: it can watch something once in a simulator and then immediately perform it in the real world.
The entire training process can even be completed with 0 gradient updates and 0 fine-tuning.
And incredibly, nobody explicitly taught it these capabilities. It “figured them out” on its own during large-scale Physical Data pretraining…

No wonder the model’s release immediately sent people online back six years:
The GPT-3 moment for robotics may really have arrived.

Physical Prompt Engineering Is Here: Robots Can Learn by Watching a 3-Second Demonstration!
I wonder if everyone still remembers.
One of the abilities that truly stunned people when GPT-3 came out was In-Context Learning.
Give the model one or several examples, and it doesn’t need to be retrained. By looking only at the context, it can temporarily figure out how to perform a new task.
What Generalist has done this time is take that entire approach and kick it straight into the “physical world”—
GEN-1.5 is a new-generation robot foundation model that the team developed through more than eight months of continuous pretraining.
The foundation remains the same: it continues to absorb large amounts of physical interaction data from the real world. But this time, a crucial new capability emerged:
Actions that just happened can also enter the Context!



Take the “cup-lid stacking” task below.
A human first demonstrates the task for 3–12 seconds. There are no gradient updates and no fine-tuning at any point. After watching this Physical Prompt, GEN-1.5 can immediately perform the task itself.
It’s basically: the moment you finish the demonstration, the little robot says, “You got it—now watch this~”
I even did a rough calculation: from learning the skill to completing the task, it seems to take only about “30 seconds”???

And the most astonishing thing about GEN-1.5 is that, in its worldview, different demonstrations can even be “stitched together” and used as one…
For example, show it one demonstration of action A, followed by one of action B, and the model can combine the two skills into a single continuous task.
It also has to figure out the transitions that nobody ever demonstrated.
Repositioning itself, adjusting its grip, changing posture, and even figuring out how to continue after making a mistake midway through—all of these can be filled in by the model itself:

And its “teacher” doesn’t even have to be a real person.
Because this model can also happily consume demos from virtual worlds… (The kid really isn’t picky.jpg)
Actions generated by scripted policies, demonstrations performed by reinforcement-learning agents, and operations carried out by humans remotely controlling robots in simulators can all be placed in its context as Physical Prompts.
The example below makes this especially clear: GEN-1.5 first watches the virtual demonstration in the simulator on the left, then directly reproduces it 1:1 in the real world on the right.
More importantly, it had never been specifically trained for this task in the simulator, nor had it practiced the task in the real world.
This is truly a case of learning online and immediately taking the job offline.

Pretty intuitive, right?
I don’t know if anyone else feels the same way, but GEN-1.5 really resembles how we use large models today.
If we want Claude to write in a new format, we usually don’t retrain a new Claude model. We simply throw examples into the Context and let the model figure it out.
With GEN-1.5, though, the logic of task completion has evolved from text continuation all the way to robotic arms. (doge)
Generalist vividly describes this capability as “Physical Prompt Engineering”—
In other words, the Prompt has gone from a sentence to a real-world action sequence.
First, demonstrate how to pull a zipper, stack cup lids, or take something out of a wallet. These actions enter the robot’s Context directly, becoming the “prompts” it uses to understand the current task~

Of course, being able to learn something in a few seconds is still a long way from being “perfectly reliable.”
Generalist tested the model on 10 tasks:
With just one Physical Prompt and 0 gradient updates, GEN-1.5 achieved an average success rate of 59%. If 5 minutes of data is provided for each task, along with 10 gradient updates, the success rate rises further to 83%.
At this stage, most of the test tasks are also relatively short and atomic operations.
The team itself acknowledges that skills learned temporarily through In-Context Learning are not yet as stable as those learned through actual Fine-tuning.
But that is precisely what has the research community so excited—
A 59% success rate may not sound that impressive, but achieving 59% with 0 training steps? That’s the scary part!!!

Nobody Specifically Taught It One-shot Learning—This Capability “Grew” on Its Own
And that’s still not the end of the remarkable things about GEN-1.5.
Because here comes the part that feels even more like GPT-3—
According to Generalist, they never specifically taught GEN-1.5 how to perform One-shot Learning!
There was no model architecture specially designed for In-Context Learning, no dedicated Meta Learning loop, and no additional objective function forcing it to learn how to improvise.
As for how these capabilities emerged, they actually appeared on their own during large-scale Physical Data pretraining… (Oh my god.jpg)

During GEN-1.5’s pretraining, Generalist initially just kept noticing that less and less data was needed for new tasks.
Hundreds of fine-tuning steps eventually became dozens; dozens of steps became 10.
Finally, with just 1 minute of data and 1 gradient Step, the robot began to learn new things.
So the team casually asked an even crazier question: What about 0 steps? It actually worked…

This also explains why so many people immediately thought of GPT-3.
Because large language models underwent a similar turning point back then—
As model sizes and pretraining continued to scale, Few-shot and In-Context Learning capabilities suddenly became pronounced enough to notice.
In Generalist’s view, robotics data may follow a similar pattern, with the key buried in the continuous real-world physical data that GEN-1.5 absorbed over the long term.
Because actions in the real world rarely happen in isolation.
After tightening the first screw, there is often a second one. After organizing one object, there is usually another. When a grasp fails, people instinctively try to grab the object again and continue the task.
In other words, data involving organization, assembly, and transportation naturally contains repetition, continuation, and recovery from failure.
When a model is pretrained for a long time on continuous trajectories like these, the actions and outcomes that just occurred naturally provide context for what comes next.
Over time, it has the opportunity to learn one crucial thing: first look at what just happened, then decide what to do next.
That is already very close to the beginnings of In-Context Learning…

I’m serious.
Let’s open our minds a little further.
If GEN-1.5’s approach can really continue to Scale, then the most important skill robots come with out of the factory may truly be just four words:
Watch! And! Learn! (doge)
References:
[1] https://generalistai.com/blog/gen-1.5#one-shot-in-context