to get there, it has to see, hear, touch and move like us. somebody has to record how people do that first.
addy crezee sat down with manish agarwal and ishank gupta, co-founders of humyn labs – the team that has filmed 1m+ first-person videos and recorded 50k+ hours of audio of how people do it. they talked about the robot age – in homes, in factories and in how we talk to machines.
what's in the conversation
- "robots as companions will be ahead of general purpose humanoids that will come and do your laundry" – the bet some of the biggest ai labs are making
- asked which sense is the harder data problem, sight or touch, both founders said the same word: "touch"
- one label off by a millimeter, times 100,000 videos, and the robot knocks the cup off the table instead of picking it up
- "very few people in the world would be using keyboard and mouse to communicate with machines. it's all going to be voice"
- "more robotic productivity than human productivity in as early as 2035"
the first robot in your home may not fold your laundry. it may keep you company.
that is not a line from a sci-fi trailer. it is the bet ishank gupta hears from some of the biggest ai labs he works with. "a couple of frontier labs actually have the companion thesis for robotics ahead of general purpose humanoids," he told us. "they believe robots as companions will be ahead of general purpose humanoids that will come and do your laundry."
but a friend has to do a lot before it can be a friend. it has to see you, hear you, understand your tone, hand you a cup without breaking it. people learn that as children. a robot has nobody to learn it from, unless someone records how people do it.
that is what humyn labs does. its team puts cameras on people's heads while they work and live: a mechanic under a car hood, a warehouse crew moving boxes, a hand watering plants with a skeleton of red dots following every joint. it records people speaking in dozens of languages. its catalog lists 1m+ first-person videos, 50k+ hours of audio and 33+ languages.
we spent half an hour with the two co-founders on what a machine has to learn before it can live next to us, and which of those lessons is the hardest.
a real tamagotchi at home
the companion idea came up in the quick-fire round, and manish agarwal answered it with a story about japan.
agarwal: "digital pets were very effective antidote to loneliness in a lot of parts of the world, especially in japan." so the next step is obvious to him: "you would have robo pets, robo companions."
addy pushed it further. why not a real physical tamagotchi at home, one that walks around the room and reacts when you say something to it?
agarwal did not oversell it. is it going to be the first thing robots do at scale? "i don't think so." two things hold it back: "affordability and technology are the two constraints." but the demand is already there. "if wishes were horses and technology was there, this could be the easiest one, because the loneliness is a reality."
gupta explained what kind of data a robot like that needs. it is not a video of one task. it is a whole day. "a data for this type of use case is what is called a day in the life of a human. so that type of data we do."
sight: machines can't read the video
to teach a robot to see like us, you would think a camera is enough. it is not.
say i film myself cutting salmon. how does that become something a robot can learn from?
agarwal: "unfortunately, machines can't read the video. there has to be something between machines and video."
to a machine a video is just pixels. someone has to write down what happened inside it. how the knife was picked up. how the fingers wrapped around it. how the wrist moved. how much force went in. that written layer is what the robot learns from.

"converting that intelligence into machine ready formats is what humyn labs actually stands for."
language models never had this problem. the internet had already written their training text. for robots, agarwal says, "everything has to be built from ground up."
"knowledge workforce, physical workforce, very, very different. for knowledge workforce you had text and code on the internet, which you could scrape. for the physical workforce, how do you get the data?"
touch: the hardest sense
the shortest answer of the interview was also the most telling.
sight or touch – which is the harder data problem right now in physical ai?
gupta: "touch."
agarwal: "touch."
there is a reason both said it without a pause. a camera can record what a hand looks like. it cannot record how hard the hand squeezes, how a surface slips, or when to let go. a friend that hands you a glass has to know all three. that part of being human is still the least recorded.
motion: one millimeter is a lot
movement is where small mistakes turn into broken things. we asked gupta whether bad data has ever caused a real failure.
gupta: i can't talk about a specific customer case, but safe to say the following holds true.
"in that one millimeter error, over a hundred thousand videos, that error compounds so crazily that the effective output you'll get in the form of action is the robot is actually hitting the cup off the table rather than picking the cup off the table."
"and by the way, one mm is a large error in robotics."
agarwal put the cost in terms anyone who has run a factory knows.
agarwal: "if you say that i'm 90% accurate, which means that out of 100 tasks, 10 tasks you have goofed up, that cascading effect of those 10 tasks on the entire production line, on your customer, on your downstream, is humongous."
"when you start operating on the interaction of physical and digital, your repercussions become real."
in a chatbot, a wrong answer is annoying. in a robot standing next to you, it is a broken cup, or worse. that is why a friend made of metal has to move more carefully than any software ever had to.
hearing: the end of the keyboard
the sense gupta cares about most is hearing, because he thinks it will become the way we talk to every machine.
gupta: "sound is going to be the interface of communication between humans and robots. i think 10, 15 years from now, very few people in the world would be using keyboard and mouse to communicate with machines. it's all going to be voice."
for that to work, a robot has to understand everyone, not only people who speak english. humyn labs built a public test for speech models called bridge. it opens with the line "ai listens to everyone except 5.5 billion people."

hearing is also harder than it looks. gupta lived in china for four years. in chinese, meaning lives in the tone, and "the same word that could mean mother could also mean horse, depending on the tone." people also switch between two languages in the middle of a sentence. "one small error in the transcription because of the code switch can teach garbage to the voice model on the other end."
a friend who mishears you is not much of a friend. a robot that hears "horse" when you said "mother" is the same problem at scale.
before the home: warehouses, factories and cars
a companion is the most personal robot, but it will not be the first one most people meet.
warehouses or homes first? gupta: "robots in warehouses will be mainstreamed before robots at home." he sees "tremendously large number of robots in the commercial world, manufacturing, mining, warehousing, maybe even delivery of food. drones are a form factor of robots." what is still missing in warehouses is the ability to make decisions on the spot. "that will evolve."
will robotics be bigger than cars? agarwal did not hesitate. "how many cars can a person have?" and then: "car covers one use case while robots are multiple use cases across different walks of life."
when do robots out-produce people? gupta: "if you believe what elon musk says, there will be more robotic productivity than human productivity in as early as 2035."
do not come for the gold rush
the last word went to agarwal, and it was a warning to anyone who reads this and wants to build a robotics startup tomorrow.
agarwal: "do not come for the gold rush. if you truly are coming in, dig deeper, anchor yourself, be agile, keep pivoting. one out of million may do that, rest will fail."
gupta's advice took one sentence. "whenever founders come to me for advice, i send them to manish."
what we took away
a robot friend is a senses problem before it is a hardware problem. it has to see, hear, touch and move like a person. each of those senses needs its own recordings of real people, and most of them do not exist yet.
touch is the gap. both founders named it in one word. sight and sound can be recorded with a camera and a microphone. how a hand feels an object is the part nobody has captured at scale.
the home comes last, and it may come through the ears. warehouses and factories get robots first. when robots do reach our homes, we will most likely talk to them, not type. that makes every language, every accent and every tone part of the job.
chapters
- 00:00 highlights
- 00:28 what humyn labs sells
- 01:01 why robot data is not like language model data
- 03:27 the salmon question
- 05:19 is $20m enough
- 09:39 a millimeter, times a hundred thousand
- 11:35 sound instead of the keyboard
- 14:46 bridge and the 5.5 billion
- 18:27 the blitz
- 21:21 a real tamagotchi at home
- 24:10 global south data
- 25:10 do not come for the gold rush
editor's note: quotes are condensed from the recorded conversation for length and clarity.
sources
- the full interview with manish agarwal and ishank gupta – thehype on youtube, sep 29, 26 min
- humyn labs – the company: training data for robotics and physical ai
- humyn labs datasets catalog – 1m+ video episodes, 50k+ audio hours, 33+ languages
- bridge report – humyn labs' public test of speech models