Neural Data Emerges as Robotics Training Bottleneck
Startups are manufacturing physical manipulation datasets from brain waves and muscle signals, betting that egocentric data collection can scale where web scraping cannot.

The Data Problem Physical AI Cannot Scrape Away
Inside a California warehouse, a worker wearing a headset pulls wooden blocks from a Jenga tower while sensors record his brain waves. This is not a neuroscience experiment. It is data manufacturing for robotics companies trying to teach neural networks how to manipulate physical objects with human-like precision.
The scene illustrates a constraint that has emerged across the physical AI sector: the sheer absence of training material. Large language models trained on trillions of tokens scraped from the open web. No equivalent corpus exists for robotic manipulation. Self-driving car companies generate their own data at enormous cost. Training from video lacks the fidelity of real-world interaction. The industry now faces a question of scale and economics that scraping cannot solve.
Encord, a London-founded data infrastructure company, has built a pilot team in San Leandro to manufacture manipulation datasets for robotics clients. The company began by helping machine-vision teams annotate existing data. As customers moved toward end-to-end learning for manipulation tasks, executives realized the data simply did not exist in usable form. Manufacturing it became the business.
Brain Activity as Model Signal
The brain wave headset comes from Zander Labs, a German neuroscience startup. Its sensors measure electrical activity in the cortex to infer mental states: intent, surprise, error detection. The hypothesis is that tagging video frames with cognitive load offers model builders a richer signal than vision alone.
Lucas Gehrke, a neuroscientist at Zander supervising the trial, argues that brain activity levels during a task reveal when a model should deploy high-effort inference versus lightweight heuristics. If a human operator shows heightened neural activity while threading an ethernet cable into a server port, that moment likely requires fine motor precision a model must learn to replicate.
Encord's trial remains exploratory. The company plans to tag an initial dataset with brain wave metadata, run it through customer models, and evaluate whether performance improves before scaling. Vineeth Velmurugan, Encord's head of robot learning and a veteran of OpenAI's robotics lab, calls this the bleeding edge of solving the data bottleneck.
At DailyTechWire, we have tracked similar experiments across Seoul and Shenzhen, where humanoid startups have begun instrumenting factory workers with wearable sensors to capture manipulation sequences. The common thread is desperation: without a YouTube-scale corpus of physical interaction data, companies are inventing new modalities to manufacture signal.
Egocentric Video and the Economics of Annotation
The robotics industry now sources training data from two primary methods. First, egocentric video collected by workers wearing head-mounted cameras, often augmented with additional camera angles and sensor streams. Second, teleoperation setups where human operators control robotic arms remotely, generating paired demonstrations of manipulation tasks.
Encord deploys both. The company collects egocentric footage from factories globally and uses its San Leandro facility to experiment with new data modalities or fine-tune specific skills. When we visited, pilots operated leader-follower rigs, paired robotic arms where one mimics the movements of the other, to generate datasets for pouring coffee (poorly, with considerable spillage) and stacking poker chips.
Storage racks held props: plastic vegetables, kitty litter scoops, fake flowers in vases, bundles of ethernet cables. These are the raw materials for training household manipulation. One pilot, Sofia Infante, maneuvered robotic arms to plug and unplug ethernet cables from server ports, the kind of work data center operators would automate if precision allowed. Operating the controls ourselves, the gap became obvious. Pincers lack the degrees of freedom and dexterity human fingers provide.
Encord also experiments with forearm sensors that detect electrical signals in muscles. Because overhead video typically does not capture the full hand, the company hopes to reconstruct a three-dimensional depiction of hand position from muscle activity, offering models a more complete understanding of manipulation geometry.
All of this footage is densely annotated with physical descriptions: "right hand tightens bolt," "left hand stabilizes object." Velmurugan estimates that densely annotated data is worth one hundred times more than raw egocentric footage for training specific tasks. It costs twenty times more to produce. On paper, that trade works. In practice, twenty times more is real expenditure, and it changes the economics fundamentally.
The Cost Barrier LLMs Never Faced
Scraping text from the open web cost frontier labs nearly nothing. Stack Overflow, Wikipedia, archived forums, digitized books: all available at the cost of bandwidth and storage. Physical training data cannot be scraped. It must be manufactured, frame by frame, task by task, with human operators wearing sensors in controlled environments.
Velmurugan estimates that breaking through will require a dataset roughly five times the size of YouTube's video corpus. That scale explains why data generation has become a distinct business rather than a research problem internal to robotics labs. The comparison between physical AI and large language models breaks down at this point. The raw material for one was abundant and free. The raw material for the other is scarce and expensive.
This is not an engineering bottleneck. It is an economic one. Model architecture continues to improve. Inference costs decline. But the cost of generating millions of hours of annotated manipulation footage remains stubbornly high, and no one has found a way to scrape it from the environment at scale.
The Vantage Point Advantage
Encord's position between multiple robotics programs offers a strategic edge. The company can observe which data techniques gain traction across the industry before any single customer sees the pattern. That visibility allows Encord to steer its own data manufacturing toward methods that show early promise in model performance.
Both pilots we observed, Infante and the Jenga operator Andrew Ceja, previously worked at Scale AI, another annotation firm. Ceja had managed a robotic trash sorter at a waste management company before joining the growing workforce building training datasets for neural networks. As the Jenga tower collapsed during one trial run, he described the work as perpetually novel, a daily exercise in solving manipulation tasks robots cannot yet handle.
Progress is being made, according to Velmurugan. With visibility into programs across the sector, he sees startups and frontier labs iterating toward what works. The dozen pilots at Encord's facility remain occupied. The question is whether this model of data manufacturing can scale to the corpus size required, and whether the economics pencil out for robotics companies already facing high capital costs in hardware and compute.
What the Neural Signal Reveals
The brain wave experiment may or may not prove useful. Tagging video frames with cognitive load is a bet that mental state offers a meaningful signal for model training. If it does, the industry will manufacture more of it. If it does not, Encord and its competitors will move on to the next modality.
What the experiment reveals is the depth of the data problem. Companies are now willing to strap electroencephalography sensors to workers pulling blocks from Jenga towers because they have exhausted easier options. Video alone does not provide enough fidelity. Teleoperation generates useful data but does not scale. Egocentric footage from factories helps but remains expensive to annotate densely.
The robotics sector is iterating toward a solution, but the solution will not resemble the data collection process that enabled large language models. It will be slower, more expensive, and more labor-intensive. That reality shapes the trajectory of physical AI and determines which companies can afford to compete at the frontier.


