Brain Waves: The New Frontier for Physical AI
- Nishadil
- July 27, 2026
- 0 Comments
- 5 minutes read
- 5 Views
- Save
- Follow Topic
Can measuring operators’ brain activity finally close the data gap in robot training?
Start‑up Encord is pairing brain‑wave headsets with its robot‑training warehouse in California, hoping that neural signals can tag data and make physical AI learn faster.
Walking into a dusty warehouse in San Leandro feels a bit like stepping onto a giant Jenga board. Wooden blocks wobble precariously, and a lone figure—who calls himself a “pilot” at Encord—carefully pulls pieces apart while a headset sits on his head. The headset does the usual: a camera that watches what his eyes see. But tucked into the band are sensors that listen to his brain waves.
Encord, a data‑tooling company, builds the pipelines that feed robots the visual and tactile information they need to understand the world. Its founders realised early on that the real bottleneck for warehouse and humanoid robots isn’t the cleverness of a neural net, but the sheer scarcity of real‑world training data. “The data simply does not exist,” admits Vineeth Velmurugan, Encord’s head of robot learning, a veteran of OpenAI’s robot lab and Berkshire Grey.
Enter Zander Labs, a German neuroscience spin‑out. Zander’s brain‑wave headset claims it can infer mental states—error, intent, surprise—directly from the operator’s cortex. By pairing those signals with the video footage of the robot‑training session, Encord hopes to create a richer, “brain‑tagged” dataset that tells a model not just what happened, but why the human acted that way.
The experiment is still in pilot mode. Velmurugan says the plan is simple: collect a modest amount of brain‑wave‑annotated video, feed it to customers’ robot‑learning models, and see whether performance jumps. If the improvement is noticeable, they’ll consider scaling the approach.
Why might a spike in the operator’s theta band, for example, matter? Lucas Gehrke, a neuroscientist at Zander overseeing the trial, explains that bursts of activity often line up with moments of decision or error. “Those spikes are clues,” he says, “that tell a learning algorithm where its attention should be highest.” In other words, the brain could act like a meta‑label, pointing out the parts of a video that are most informative.
Encord’s broader data strategy leans heavily on what the industry calls “egocentric” video—cameras mounted on workers that capture the world from a first‑person viewpoint. The company aggregates such footage from factories across the globe, then enriches it with extra angles, sensor streams, and now, neural data. At the San Leandro site, pilots use leader‑follower rigs: a human‑controlled arm moves in sync with a twin robot arm that records the motion for later replay.
Typical tasks range from the mundane to the maddeningly precise: pouring coffee from a pot (the sloshing liquid is a nightmare for vision systems), stacking poker chips, or plugging Ethernet cables into a server rack. The latter, demonstrated by Sofia Infante, highlights a persistent gap—robotic pincers still lack the dexterity of human fingers, making even simple plug‑in tasks a stretch goal.
Beyond brain waves, Encord is experimenting with another physiological signal: electromyography (EMG) sensors strapped to the forearm. By reading muscle activation, the team hopes to reconstruct a 3‑D model of the hand’s pose, even when the camera can’t see the fingers. Combined with dense textual annotations—e.g., “right hand tightens bolt”—these multimodal streams could give large‑language‑model‑based controllers a much clearer narrative of what’s happening.
Cost is the ever‑present specter. Velmurugan estimates that richly annotated, multimodal data is about 100 times more valuable than raw egocentric footage, yet it costs roughly 20 times more to produce. “Twenty times more” is still a big number when you’re trying to train fleets of warehouse robots, but it’s a trade‑off that makes sense if it cuts the time to market for new robotic capabilities.
The bigger picture mirrors the early days of large‑language models, where developers scraped the entire web for text. For physical AI, the “web” is the warehouse floor, the factory line, and now, the operator’s brain. As Velmurugan puts it, “If we need a dataset five times the size of YouTube’s video corpus to break through, we have to start thinking about what else we can capture besides pixels.”
Whether brain waves become a standard tag in the robot‑training pipeline remains to be seen. For now, Encord and Zander are collecting the first handful of neural‑rich clips, hoping the data will prove that a flicker of thought can indeed guide a machine toward better, safer, and more adaptable physical intelligence.
Editorial note: Nishadil may use AI assistance for news drafting and formatting. Readers can report issues from this page, and material corrections are reviewed under our editorial standards.