Embodied AI data collection has a new path.
Before we formally introduce it, let’s start with a quick interaction: leaving AI aside and without looking anything up, can you name the existing approaches to embodied AI data collection within five seconds?
Teleoperation data collection with real robots, body-free wearable collection (UMI/Ego), simulated synthetic data, and distillation from online videos are the four mainstream approaches.
(In no particular order—just a quick list.)
If we break them down further, we could come up with a much longer list.
At least, that’s what AI gave me as an answer. Take a look—

Recently, a new approach has begun to emerge and gain momentum. It shifts the focus of data collection beyond externally observable actions and further toward the body’s motor output.
This is the much-discussed neural electromyography (sEMG, Surface Electromyography) approach.
OriginFlow is one of the most closely watched startups pursuing this route.
Founder and CEO Qin Shentao, born in 2001, is currently pursuing a PhD at Tsinghua University’s School of Vehicle and Mobility.
In August last year, he founded OriginFlow in Beijing. As of May this year, the company had publicly disclosed the completion of three funding rounds—an angel round, a strategic round, and a Pre-A1 round—with total funding exceeding RMB 500 million. Investors include BlueRun Ventures, Oasis Capital, and Monolith, among others.
Qin Shentao told us:
Between real-world industrial scenarios and models, there is a missing layer capable of abstracting unstructured, high-precision, and highly detailed physical interaction information into trainable representations.
That is why he chose to start a company and build infrastructure for Physical AGI, with the hope of making neural electromyography one of the entry points for embodied AI data collection.

Qin Shentao said that, in the short term, NeuroScale would serve more as an incremental enhancement to approaches such as UMI and Ego-centric first-person vision. However, invisible, unobtrusive data collection is certain to become an important trend in the future.
When new changes emerge, the more appropriate strategy is to evolve collaboratively with existing technologies.
Collecting Embodied Data Through a “Neural Interface”
Let’s start with everyday life and get a feel for what OriginFlow is currently working on.
When a person tightens a screw, steadily holds a bowl of soup, or grips a soft object, the body continuously provides subtle feedback based on the circumstances.
When do the fingers begin to tighten? How should the force increase or decrease? … In reality, people do not consciously think through these questions one by one. Instead, in an extremely short period of time, the body completes and executes a series of judgments concerning contact, friction, weight, and stability.
OriginFlow hopes to record this layer of response as part of the data used by embodied AI.

To understand why the company is betting on this direction, we need to go back to Qin Shentao’s undergraduate years.
During his undergraduate studies, Qin led his team to win nearly every robotics competition they could enter. It was also during this period that he began to think more deeply:
Keyboards, mice, and voice are not the most direct ways for humans to communicate with machines.
Human hands are the most dexterous tools. If machines could understand the true intentions behind hand movements, human–machine interaction could reach a whole new level.
That was when neural interfaces entered the field of vision of Qin Shentao, who has long been “obsessed” with human–machine integration.
Neural interfaces collect surface electrical signals associated with the peripheral motor nerves—that is, neural electromyography signals. This allows them to avoid some of the challenges faced by brain–computer interfaces, including high invasiveness, large channel counts, and the need to map signals to specific brain regions.
Since the emergence of neural electromyography, the technology was long confined to university laboratories.
Later, as wireless sensing, dry electrodes, and AI algorithms continued to evolve, neural electromyography began to be used as a sensing entry point for human motor intention. It has since found broad applications in biofeedback training in rehabilitation medicine, athletic movement analysis, occupational ergonomics and fatigue assessment, and discrete-action control of myoelectric prostheses.

△ Image generated by AI
Here is a typical example.
In 2019, Meta’s core R&D division, Reality Labs, acquired CTRL-Labs, a noninvasive neural interface company.
Before the acquisition, CTRL-Labs had already released a developer kit that included a main device about the size of a watch and a tethered component containing electrodes.
Using a 16-channel electromyography sensor, it monitored muscle electrical signals at the wrist transmitted from motor neurons to the fingers. AI algorithms then decoded these signals into digital commands such as clicks, swipes, and gestures.

Two years later, Meta publicly demonstrated the Meta Neural Band, a neural-interface wristband prototype based on CTRL-Labs’ technology. The demonstration showcased its potential to control AR interfaces by capturing tiny neural signals, enabling actions such as typing in midair and selecting menus.
Zuckerberg said at the time that, because this technology could precisely sense a user’s intention to perform an action—even before the action had actually been made—this interaction method offered “almost infinite control.”
Last year, Meta’s third-generation smart glasses introduced a neural-input wristband, allowing users to control the interface on the glasses directly through hand movements, such as imagining a swipe or click.
(An aside: The “pinch” gesture used by Apple Watch for confirmation, closing windows, and similar functions is primarily based on optical sensors.)

While Meta is focused on making it easier for people to click, swipe, or enter text, Qin Shentao’s thinking centers on what Physical AGI currently lacks.
In his view, the severe shortage of high-quality physical interaction data has become a key constraint on the advancement of embodied AI robots.
Each existing approach to embodied AI data collection has its own drawbacks. Could neural electromyography serve as an entry point for further supplementation?
Moreover, collecting neural electromyography signals might allow systems to bypass differences in the materials and sensors used at specific points of contact. Starting from the respective actuation sides of humans and machines, it may be possible to identify a shared action representation that connects muscle activation, tendon force, and joint movement.
Collection Starts with a Wristband
In 2025, Qin Shentao, then a PhD student, officially started his company and incorporated OriginFlow. Its vision is to build an “action foundation” beyond text and video for Physical AGI.
Not long afterward, the team introduced the NeuroScale technology stack.
NeuroScale is not simply a matter of adding an EMG wristband to existing embodied AI data-collection equipment.
It is a data and model system spanning signal acquisition, physical quantity reconstruction, action representation, and cross-embodiment transfer.
The system uses a noninvasive neural motor interface as its signal entry point, combining neural electromyography, first-person vision, and multimodal information such as IMU data. Through the PULSE foundation model, it reconstructs the posture, contact forces, and driving forces involved in human manipulation, and converts each real-world operation into machine-learnable Human Tokens.

NeuroScale focuses on two core questions over the long term.
First, how can Human Data Scale Up be achieved?
In other words, how can real-world operations be continuously recorded and more high-quality physical interaction data produced while minimizing interference with people’s natural perception and movement?
Second, how can Cross-Embodiment Transfer between humans and robots be achieved?
The goal is to allow human action experience, after being represented and adapted, to be transferred to robots with different structures.
Underlying these two questions is an even more fundamental technical problem: In what form should human actions enter a foundation model?
Text already has a relatively mature Token representation, and vision is gradually converging on Patch or Latent representations. Yet the action modality still lacks a widely accepted standard representation.
OriginFlow decomposes a physical operation into three interrelated spaces:
-
Motion Space, which describes hand posture and movement trajectories
-
Tactile Space, which describes normal force, tangential force, and contact feedback
-
Tendon Space, which describes muscle and tendon driving forces and joint torque
Driving forces produce movement, movement creates contact, and contact ultimately generates applied force. Together, the three form a physical chain of causality.
The Human Tokens described by OriginFlow are precisely an action representation built on these three types of physical quantities.
“Formally speaking, a human can be viewed as a special case within the configuration space of embodiments. Therefore, the transfer from humans to robots is essentially a subproblem of cross-embodiment transfer.” Qin Shentao said.
More specifically, large-scale Human Data is responsible for covering as broad a distribution of human actions and skills as possible. Cross-Embodiment Data, which is relatively smaller in scale but encompasses a sufficiently diverse range of robot configurations, provides alignment anchors between humans and different robotic embodiments.
By combining the two types of data, the model may learn a shared action representation that is relatively decoupled from any specific embodiment. The same human movement and exertion information can then be redirected to robotic embodiments with different degrees of freedom and actuation methods.
Therefore, OriginFlow is not trying to copy human manipulation data directly to robots. Instead, it first seeks the actions, driving forces, and contact relationships that can be shared across different embodiments, and then performs adaptation and mapping.
In the real world, NeuroScale starts with a wristband.

The wristband is called OriginKit Gen 1.0. It is designed to be slightly wider than an Apple Watch band, while being smaller and lighter overall.
The wristband collects the wearer’s surface EMG signals at the microvolt level. The system uses 16-channel acquisition, with an information data rate of approximately 96 KB per second, and continuously models hand movements.

However, NeuroScale’s deployment does not rely on the wristband alone.
The wristband by itself can only detect changes in forearm muscle activity. It is difficult for the system to determine exactly what those changes mean the fingers are doing.
The signals collected by OriginKit Gen 1.0 enter the NeuroScale data pipeline together with first-person vision, IMU data, and other information.
After the raw waveforms are aligned, calibrated, and processed, the team’s proprietary foundation model, PULSE, extracts clues related to hand posture, movement trajectories, contact forces, and tendon exertion. These correspond to the three core physical quantities of Motion Space, Tactile Space, and Tendon Space.
The hardware handles filtering, differential processing, and motion-artifact suppression, while the model handles neural-signal encoding and strongly supervised learning, gradually organizing the data into action representations that machines can learn.
These are what the team calls Human Tokens.
During WAIC, Qin Shentao and his team demonstrated their current results: a demo of PULSE 0.2:
The wristband collects 16-channel surface EMG signals. PULSE uses these signals as input to continuously model hand movements and observe changes in exertion during a finger-pinch motion.
Unlike discrete gesture recognition, PULSE focuses on continuous hand tracking and fingertip-force representation.
When the user performs a finger-pinch motion, the system can observe changes in exertion in real time.
However, reconstructing movement from a segment of neural signals does not mean that the resulting data can already be used directly to train robots.
The raw data must still go through multi-device clock synchronization, individual calibration, motion-artifact processing, task segmentation, action and force annotation, quality filtering, and cross-embodiment mapping.
To support this, OriginFlow has also built a Data Infra covering data production, processing, evaluation, and training.
Among its components, ORACLE uses multimodal models to automatically annotate actions, forces, semantics, and task segments. The unified all-modality foundation, CHORD, handles temporal and representational alignment across neural electromyography, vision, IMU, language, and robot states.
The data then passes through a series of quality checks, including physical validity, task validity, annotation validity, model value, and real-robot evaluation. Only data that generates tangible value in model training or robot tasks enters the final deliverable dataset.
In other words, it is not enough to collect the data. The system must also be able to process massive volumes of raw signals into trainable, evaluable, and reusable data assets at sufficiently low marginal cost.
From Human Data to Enhance Human
The PULSE 0.2 demonstration represents only a small part of the company’s business, and the model itself is merely a milestone along the way.
PULSE 0.3, currently under development, is exploring the relationship between tendon exertion and the forward kinematics of the hand.
The wristband mentioned earlier is also only the current collection terminal, not the final product form.
For an approach that depends on long-term, large-scale data collection, whether the device can be naturally accepted by people determines whether data can continue to be generated.
Qin Shentao said the goal is to make people willing to put on the device in the first place—and willing to wear it for long enough.
At its core, OriginFlow hopes to continuously record people’s actions in the real world while changing their existing patterns of perception and movement as little as possible, and then process that information into Physical Tokens.
Only then can data collection move beyond data-collection factories and specially constructed workstations and enter everyday life and real production workflows.
At that point, the “trillion hours of data” currently considered out of reach in embodied AI would require recording only around ten days of behavior from the entire human population.
Qin Shentao describes this process as the continuous scaling-up of data.
Looking ahead, we will drive the scale of multimodal Human Data from tens of millions of hours and hundreds of millions of hours gradually toward trillions of hours.
The team’s vision for the next generation of physical interfaces is to move toward forms that are lighter, more natural, and closer to consumer products.
In their vision, people will eventually need only to wear consumer-oriented glasses, watches, or even lighter wristbands to continuously model themselves in everyday life and gain convenience from their personal models.

△ Image generated by AI
Overall, OriginFlow’s current work remains in the first stage of its development roadmap.
The company has laid out a three-stage development path.
Stage One: From Human.
The first step is to collect human actions as extensively as possible, systematically distilling and converting people’s operational experience in the physical world into machine-learnable data and representations.
OriginFlow hopes to fully reconstruct human action output without interfering with people’s natural perception and movement.
The entire collection solution follows a noninvasive principle, ensuring that people can maintain their original state while their movements are accurately observed and reconstructed by the system.
Stage Two: With Human.
At this stage, the company will build a new generation of AI hardware, including robots, and integrate it into everyday work and life in an Always-on form, allowing more people to benefit from the next generation of human–machine Interfaces.
Stage Three: Enhance Human, aimed at the next 10 to 30 years.
Of course, a long process of engineering validation still lies between a single finger-pinch demonstration in PULSE 0.2 and the trillion-hour scale of multimodal Human Data.
One More Thing
It is reported that other Chinese startups, including BrainCo, Octopus Dynamics, SnowOrigin, Nianxiang Technology, Unbounded Neural, and Shouyi Technology, are also using neural electromyography to collect embodied AI data.