AI

Are brain waves the next unlock for physical AI?

The frontier of physical AI is a Jenga game in a warehouse in San Leandro, California.

That warehouse is occupied by Encorda company that builds data tools used to train AI models. Andrew Ceja is a pilot – the company’s term for its robot trainers – and he carefully pulls wooden blocks from a rickety tower while wearing headphones with a camera that tracks what he sees. That alone is quite common for collecting robot training data, but this headset contains sensors that measure his brain waves as he carefully disassembles the block tower.

Encord is one of a small but growing number of startups betting that the next real limitation for humanoid and warehouse robotics won’t be model architecture, but instead the sheer scarcity of real-world physical training data. Instead of just helping robotics companies manage the data they have, Encord is building a business around producing the data they don’t have.

The brainwave headset that Ceja wears was built by Zander Labsa German neuroscience startup betting on measuring brain activity – to infer mental states such as error, intention and surprise – could create a more useful data set to train models. Encord’s work with Zander is currently a pilot project; Encord says the goal is to build an initial data set of brain waves, run it through customer robot models and evaluate whether it actually improves performance before deciding whether to scale it up.

Lucas Gehrke, a Zander neuroscientist overseeing the work, says the amount of brain activity used at any given time during a given task offers clues to modelers trying to figure out when to deploy their models with the highest effort.

According to Vineeth Velmurugan, head of robot learning at Encord, this is the “blood point” of efforts to solve the robotics data bottleneck. Velmurugan, a veteran of OpenAI’s robotics lab and warehouse automation company Berkshire Gray, joined Encord to build the company’s internal data creation team.

See also  Top OpenAI, Google Brain researchers set off a $300M VC frenzy for their startup Periodic Labs 

Encord was founded to help companies building machine vision applications annotate data and evaluate models. When their customers (Velmurugan says they work with many leading robotics companies, but he’s not allowed to name them) started applying end-to-end learning to robot manipulation tasks, executives realized they had to produce training data themselves, rather than simply manage it. “The data simply does not exist,” Velmurugan said.

The bet that generative AI can do for robots what it has done for chatbots continues to hit the same wall. LLMs are built on text from all over the internet, and more. Finding the same raw materials to teach neural networks about physical manipulation is a challenge: self-driving car companies collect it themselves, but that is difficult to scale. Training via video can work, but it lacks the reliability of real-world data. Velmurugan says it will take a data set about five times the size of YouTube’s video corpus to break through – a scale that helps explain why data generation itself has become a business and not just a research problem.

Feed your self-centered data needs

Companies building robot brains are now turning to two main sources: “self-centered” video collected by workers wearing cameras, often supplemented with additional camera angles and other metrics, and data collection from remotely operated robots. Encord does both: it collects egocentric data from various factories around the world and uses its facility in San Leandro to experiment with new modalities, such as brain waves, or to collect data sets around specific skills to refine them.

See also  Intuit learned to build AI agents for finance the hard way: Trust lost in buckets, earned back in spoonfuls

When TechCrunch visited, pilots used leader-follower rigs—linked robotic arms, one controlled directly by a human operator and one that mimics its movements—to create data on tasks like pouring coffee from a pot into mugs (very sloshing) and stacking poker chips. “Every humanoid company has asked us for these pieces,” Velmurugan says.

Storage racks held boxes of fake flowers in vases, books, plastic vegetables, litter boxes and scoops, bags and bundles of wire, the stock in trade for training manipulators for household tasks.

At one of these stations, another pilot, Sofia Infante, maneuvers robotic arms to connect and disconnect Ethernet cables from the back of a server—the kind of work that data center operators would like to automate, if only robots could manipulate them with the precision required. As I turned the controls, I understood why that’s still out of reach: pliers are far less dexterous than human fingers and lack the degrees of freedom we take for granted in our arms.

Another new data modality Encord is developing uses a series of sensors attached to the forearm to detect electrical signals in the muscles. Video taken of human hands manipulating objects typically doesn’t capture the entire hand, but Velmurugan hopes to be able to create a 3D representation of where the hand is at any given moment based on the arm sensors, creating a more robust understanding for models.

Encord’s datasets are annotated with physical descriptions of what each video contains (“right hand tightens bolt”) to help LLM-based models understand what’s happening. Velmurugan estimates that this kind of dense annotation is worth a hundred times as much as “junky ego data” for training specific tasks, and costs only twenty times as much to produce on paper, which is a good trade.

See also  Tana snaps up $25M as its AI-powered knowledge graph for work racks up a 160k+ waitlist

But “20 times more” is still real money, and that’s the catch: scraping text from the Internet, the way LLM makers built their models using Stack Overflow and the rest of the Internet, cost frontier labs virtually nothing. Generating physical training data does not, and that is the limit of the physical-AI-as-LLM equation. This kind of data needs to be manufactured and not just collected, and that changes the economics of building these models.

Velmurugan says progress is being made. Encord’s visibility into programs across the industry allows him to see both startups and frontier labs figuring out what works and what doesn’t to improve physical AI models. That vantage point – located between many robotics companies at the same time – is also part of Encord’s pitch. It can discover which data techniques are gaining traction across the industry before a single customer can.

That will keep the dozen or so pilots at Encord’s facility busy. Both Infante and Ceja are part of a growing workforce developing the building blocks for neural networks; They previously worked at Scale, another AI data annotation company, before joining Encord.

Ceja had worked at a waste management company, where his interest in technology made him responsible for keeping a robotic waste sorter in good working order. Now that the Jenga tower is falling, he says he’s enjoying the challenge of solving training tasks for robots: “It’s something new every day!”

When you make a purchase through links in our articles, we may earn a small commission. This does not affect our editorial independence.

Source link

Back to top button