Skip to content
Main Site News Console

Silicon Valley’s Hottest Embodied Model: Learns from a Single Demonstration—No Post-Training Required

· 量子位
国内AI

Silicon Valley’s Hottest Embodied Model Yet! No Post-Training Required—Learn It After Watching Once

Embodied intelligence is moving toward its “GPT moment.”

A major breakthrough in embodied intelligence is coming!!!

Just last week, Generalist released Gen 1.5, an embodied brain capable of learning actions lasting 3 to 12 seconds. Now, North American embodied AI startup Skild AI has released a new robot foundation model, S1, extending the duration of robot in-context learning tasks to more than 10 minutes.

This time, S1 focuses on in-context learning (ICL):

There is no need to specifically learn new operation data during post-training. The robot only needs to watch a video of a human demonstration and can then “monkey-see, monkey-do,” directly completing an entire series of complex operations it has never encountered before.

In official demonstrations, the robot completed long-horizon tasks including making pancakes, brewing coffee, repotting plants, and assembling equipment. On unseen tasks, it achieved a success rate of 66%—far exceeding that of language-prompted VLA, which managed only 9%.

Moreover, even using the traditional post-training fine-tuning approach, matching the performance S1 achieves after watching a demonstration just once would require approximately 380 task demonstrations.

Incredible!

It is fair to say that this recent wave of work focused on in-context learning has reignited many people’s hopes for embodied intelligence.

One Twitter user said that general-purpose robots might appear in two years, rather than seven.

Another user said that from Rhoda, to last week’s Generalist, and now the 10-minute tasks demonstrated by Skild S1, the robot’s “GPT moment” seems to be approaching.

Because once this capability truly generalizes, developing robot skills in the future may genuinely become as simple as writing a Prompt for ChatGPT.

Is it really that remarkable? Let’s take a look.

From the BERT Era to the GPT Era

To understand how S1 performs in-context learning, we first need to look at how traditional robots learn new skills.

In the traditional pipeline, the robot’s learning process is actually quite similar to that of a large model:

First, it undergoes pre-training on massive amounts of data to acquire basic capabilities. It then enters a specific setting, collects data for a particular task, and learns specialized skills through post-training.

The difference is that although robots and large models follow similar processes, their data costs are completely different.

Large models obtain much of their pre-training data from the internet. Even in specialized areas such as programming and Agents, much of the post-training data can be quickly obtained online or even generated automatically.

Robots, however, face a much greater challenge: both pre-training data and post-training data must be collected in the real world.

That is why Skild AI stated directly in its technical blog:

If a robot foundation model still requires dozens or hundreds of hours of real-robot data and another round of post-training every time it learns a new task, then I have to ask: where exactly is the “foundation” in this “foundation model”?

They even cited existing research indicating that if enough data is available for a new task, a model trained from scratch may be able to match a foundation model that has undergone pre-training and fine-tuning.

So what exactly is the point of embodied pre-training?

Skild’s answer is: in-context learning.

To provide a point of reference, Skild uses the development path of large models as a comparison.

Early BERT models were already very powerful, but whenever they encountered a new task, it was often still necessary to prepare new data and perform another round of fine-tuning.

What truly changed the game was the in-context learning capability that gradually emerged after GPT-3:

Without modifying the model weights, simply providing a few examples in the Prompt allows the model to temporarily “learn” a new task.

Based on this, Skild believes that robots are still stuck in something resembling the BERT era. What they truly want to achieve is the same paradigm shift for robots.

In other words, the real value of pre-training should not merely be to reduce the amount of data collected during post-training. It should ultimately give robots the ability to “learn directly from context.”

In-Context Learning

So what exactly did S1 learn from context this time?

Skild believes that a robot’s in-context learning capability can truly be evaluated along two dimensions:

First, whether it can learn new skills that never appeared in the training data; and second, whether it can recombine existing skills to complete a long and complex new task.

For example, the robot only needs to watch a video of someone flipping pancakes to learn how to make pancakes—even if this skill has never appeared in its training data before.

At the same time, compared with the second-long videos demonstrated by Gen-1.5 last week, S1 extends the duration of in-context learning to as long as 10 minutes.

Tasks such as repotting plants require dozens of operations to be completed consecutively. In demonstrations of these long-horizon tasks, the robot must not only imitate what it sees, but also understand:

Which step it is currently on, what to do next, how to combine different skills, and how to continue after making a mistake midway through.

In other words, the robot must truly understand the intentions behind the tasks and actions in the video demonstration, adapting its behavior to each situation rather than merely performing behavior cloning.

In the plant-repotting task, it took only 11 minutes from the start of recording the demonstration to the robot attempting the task itself, surpassing traditional post-training in terms of efficiency.

Finally, throughout the entire process, S1 used neither fine-tuning nor post-training. The model weights remained completely unchanged, and the same set of weights was used to complete every task presented in the blog.

This essentially parallels the transition in large models from BERT, which required fine-tuning for specific scenarios, to GPT-3, which could complete OOD tasks after simply watching demonstrations.

The only difference is that the Prompt is no longer language-based. Instead, it consists of action videos that can better align with downstream tasks.

Experiment: Is ICL More Scalable?

In the experiments, Skild first compared the performance of ICL video Prompting with that of traditional language-prompted VLA on tasks that had appeared in the training data and tasks that had not.

The results showed that when the training data amounted to only 1,000 hours, language Prompting performed better. As the amount of data increased, however, ICL began to pull ahead.

When the data scale reached 100,000 hours, ICL achieved a success rate of 96% on tasks seen during training, compared with 89% for language Prompting.

On the truly critical OOD tasks, ICL achieved a success rate of 66%, while language Prompting managed only 9%—a gap of more than sevenfold.

To verify S1’s generalization ability, Skild conducted additional tests under different experimental conditions.

The results showed that the performance degradation of language-prompted VLA was as much as three times greater than that of ICL.

In other words, ICL does not learn to mechanically repeat actions in a fixed environment. Instead, it can replan its behavior based on the current environment.

Finally, in terms of one-shot learning, S1 achieved a 66% success rate on new tasks without any post-training and after watching only a single video demonstration.

By contrast, traditional VLA required post-training on approximately 380 demonstrations to reach the same level.

Of course, after increasing the number of demonstrations to 2,000, traditional post-training ultimately reached 86%.

So the conclusion may not be that “robots will no longer need training,” but rather:

A new skill that once required hundreds of repetitions to teach may now be performed at a passing level after the robot watches just one demonstration.

This is also what Skild calls the ICL scaling law:

The more data there is, the more the model learns—not only more skills, but also how to become increasingly proficient at “learning skills from demonstrations.”

Who Is Skild AI?

Skild was founded in 2023 by two familiar names in CMU’s robotics community: Deepak Pathak and Abhinav Gupta.

Both are professors at Carnegie Mellon University’s Robotics Institute. Deepak primarily researches robot learning, reinforcement learning, and computer vision, while Abhinav is an expert in computer vision and self-supervised learning.

As early as 2022, the two collaborated on WHIRL, which enabled robots to perform one-shot imitation by watching human videos. In that sense, S1’s approach did not come out of nowhere.

As a prominent company in North America’s embodied intelligence sector, Skild raised $300 million in Series A funding in 2024 shortly after emerging from stealth, reaching a valuation of $1.5 billion.

In January this year, the company completed a $1.4 billion Series C round, bringing its valuation to more than $14 billion. The round was led by SoftBank, with NVIDIA, Bezos, and others continuing to participate. In just over two years, the company’s valuation had increased nearly tenfold.

Compared with its peers, Skild occupies a particularly distinctive position in North America’s embodied intelligence landscape.

PI is more focused on building a “GPT for robots”; Generalist recently emphasized the concept of a one-shot learner; and Genesis AI and Sunday have demonstrated ambitions to develop full-stack systems. The phrase Skild has consistently emphasized is: Any robot, any task, one brain.

The company places greater emphasis on cross-embodiment capabilities, aiming for the same Skild Brain to run on different bodies, including robotic arms, quadruped robots, and humanoid robots.

Put simply, it wants to become Android for the age of robotics: placing one intelligence layer inside many different bodies.

This also determines Skild’s data strategy: use everything.

Skild categorizes robot data along three dimensions: hardware proximity, diversity, and scalability:

Real-robot teleoperation is closest to the hardware but expensive. First-person human videos are the easiest to scale but are more distant from the robot’s body. Simulation is inexpensive and can generate large amounts of data, but it introduces a sim-to-real gap.

Therefore, for Robot Teleop, UMI, Egocentric Video, and Simulation, Skild has essentially adopted an all-of-the-above strategy.

S1’s ICL is also a natural extension of this approach:

In September 2025, LocoFormer was released. In February 2026, Skild achieved In-Domain ICL for the first time. In May, it demonstrated pancake flipping for the first time. In August, S1 was released, pushing in-context learning directly into OOD long-horizon tasks lasting up to 10 minutes.

So the question truly worth关注 now may no longer be whether robots can learn more skills, but rather:

Can the cost of teaching a robot a new skill really be reduced from “teaching it hundreds of times” to “you do it once, and it watches once”?

If this scaling law continues to hold, then the so-called GPT moment for robots may genuinely be more than just another marketing buzzword.

References:

[1] https://www.skild.ai/blogs/s1