From Movement to Intent: A New Training Layer for Physical AI

Image credit : TechCrunch
Robots are becoming better at copying human movement. They can watch a person stack objects, tighten a bolt or guide a robotic arm through a warehouse task.
But movement alone leaves out an important part of the lesson.
A robot can observe that a person changed direction. It cannot easily tell whether the change came from surprise, hesitation, rising effort or the realization that something had gone wrong.
That is the gap Encord and German neuroscience startup Zander Labs are exploring. In an early trial, human trainers wear headsets that record first-person video and brain-wave activity while completing physical tasks. The goal is to test whether cognitive signals can add richer context to the data used to train robotics models.
A Falling Jenga Tower Can Reveal More Than Motion
Inside Encord’s robotics-data facility in San Leandro, California, a trainer carefully removes blocks from an unstable Jenga tower.
A camera captures what he sees. Sensors in the headset record neural activity as the tower becomes harder to manage.
The visible footage shows his hands, the blocks and the final correction. The brain-wave data may indicate when he notices instability, increases concentration or recognizes an error before the tower moves.
That distinction could help a robotics model understand not only what action happened, but when the task became difficult and why a new response was needed.
This is not an attempt to read private thoughts. Zander Labs is focused on broader cognitive states such as intent, surprise, workload and error recognition—signals that may help label important moments inside a physical task.
Physical AI Cannot Learn From the Internet Alone
Large language models benefited from an enormous supply of existing text. Physical AI does not have an equivalent archive of carefully recorded real-world behavior.
Robotics data often has to be produced deliberately. Human trainers wear cameras, operate paired robotic arms and repeat tasks under controlled conditions. Each session must then be synchronized, described and prepared for model training.
Encord is expanding from software for managing and annotating vision data into the more demanding business of creating the missing data itself. Its teams collect first-person footage, teleoperation data and sensor information from factories and dedicated training environments.
The company’s head of robot learning, Vineeth Velmurugan, argues that the constraint facing robotics may be less about model architecture and more about the shortage of high-quality physical training examples.
Unlike web text, this data is expensive to manufacture.
Brain Signals Could Help Find the Moments That Matter
Not every second of a demonstration is equally useful.
A robot trainer may perform a task smoothly for several minutes before encountering one difficult object, unstable grip or unexpected result. Those few seconds may contain the most valuable learning opportunity in the entire recording.
Brain-wave signals could help researchers locate those moments faster.
If neural activity suggests increased effort or error recognition, model builders may be able to identify where a task demands stronger reasoning or a more capable model. The information could also help distinguish ordinary movement from moments that require adaptation.
Encord is exploring other data sources for the same reason. Sensors attached to the forearm can capture muscle signals that video may miss, potentially helping reconstruct hand position and force more accurately.
Better Data May Matter More Than a Bigger Robot Model
The robotics industry has spent years improving hardware, foundation models and simulation.
Yet physical intelligence still breaks down on tasks humans consider routine: plugging in a cable, handling a flexible object, pouring liquid without spilling or recovering when an object shifts unexpectedly.
These failures often come from missing context rather than missing motion.
A robot may know where the hand moved but not how much force was required, whether the operator expected resistance or when the human decided the original plan was failing.
Combining video, muscle activity and cognitive signals could create a more complete record of how people solve physical problems—not just how their bodies move through them.
The Experiment Still Has to Prove Its Value
The collaboration remains a trial.
Encord plans to build an initial brain-wave-tagged dataset, test it with customer robotics models and measure whether it produces meaningful performance improvements before expanding the work. No public results yet show that EEG-enhanced training consistently makes robots more capable or reliable.
There are practical limits as well. Brain signals are noisy, vary between individuals and introduce serious questions around consent, storage and appropriate use.
Even so, the experiment points toward a more ambitious form of robot learning.
The next generation of physical AI may not learn only by watching what humans do. It may also learn from the moments when humans hesitate, concentrate, detect failure and decide to try something different.
Source : TechCrunch



