Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
Words
5,279
Runtime
28:18
Speaking pace
187wpm
Reading time
22min
187 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
[music] I think uh all of you know here that how much progress has AI made in the last five years. Uh seems like it's more than uh last 100 years of progress. And what is happening especially in AI is AI came for language then speech audio video and everybody's excitement is next is robotics okay and this excitement has is going on peak uh this year uh for some reason which is still beyond my uh understanding and you can also see uh leaders like Jensen talking about physical AI as the
94 words, the words spoken in the first 30 seconds at 187 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 302 |
| Average words per sentence | 17.5 |
| Longest sentence | 215 words |
| Questions asked | 18 |
| Sentences containing a number | 29 |
Most used terms
Filler phrases
203 in total: uh 135 · like 40 · kind of 10 · you know 8 · actually 4 · right? 3 · I mean 1 · sort of 1 · um 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
[music] I think uh all of you know here that how much progress has AI made in the last five years. Uh seems like it's more than uh last 100 years of progress. And what is happening especially in AI is AI came for language then speech audio video and everybody's excitement is next is robotics okay and this excitement has is going on peak uh this year uh for some reason which is still beyond my uh understanding and you can also see uh leaders like Jensen talking about physical AI as the next frontier of robotics.
So it seems like if you open Twitter or LinkedIn, it seems like robots robotics is already here and we'll have home robots in our homes sometime by December uh as uh as it's being said by multiple people. But I'm here to highlight that robotics has been almost here for the last 70 years [sighs] now. Those of you who got into robotics in the last four or five years, go and take any any any course in robotics and if you can distinguish the videos in robotics from 30 years ago versus today, uh I owe you a dinner.
So to give you an example, let's look at this robot. So this robot here is an image and the goal of this robot is to look at this image of blocks and arrange these blocks in front of it so that it looks the same pattern as in the image. Okay, this is an extremely simple task. But if you think about this, it's a 3D block. You have to look at understand 3D from each side. And the robot can do this pretty precisely. Any guesses how old this result is?
Anyone? It's a guess. You can make any guess. Huh? >> 40 years. Okay, that's the highest. [panting] This is from 1960s. This is way before there were computers. Uh, okay. So, this is this is called MIT copy demo. Uh, this was the start of uh this is this predates AI. Uh, as you know, today look at this one. >> One exhibit at the Nuclear Congress in Philadelphia has the perfect formula for >> what's happening here. The guy behind the scene is controlling these robots, a leader follower system.
And if you look at any big lab, any big company, any big academic lab, how they get data, they use teley operation system, which follows the exact same principles. Any guesses for the year for this one? >> Now, everybody would correct aggressively on the year, but this is from 1957, 68 years ago. Okay? Ask anybody in your family who is older than actually they may not be around like so because this is for 60 you have to ask somebody who is 80 year old what was the state of technology at that time like RAM for few KB of RAM you would have a size of a room uh like this big of a setup and this is from that time when people could make it work on not just this you keep coming back every decade the videos look as cool like here is robot a humanoid juggling playing foosball uh with human and this is also from you can see this is all they were all predate deep learning this from 8 years ago so this robot is a result from Berkeley can clean up the whole table can arrange items in a in a in a box this was done by one grad student with a single GPU machine nothing more than that okay now this picture of and if I colorize all these videos from past and apply genai filter for the modern voice you cannot tell that this [laughter] video is from 19657 or this is from today And these are not even the oldest results the oldest ones go to 1940s.
Okay. So what is this that if you look at last 70 years tech every other technology every other technology has come way far language understanding computer vision mobile phone chips why is robotics is stuck in this primitive age uh for for this long and the reason is uh a general brain. Robotics has always been approached as a hardware problem from ground up and this the reason that everything around robotics has progressed but robotics is still stuck in the same land from the last 70 years.
So there's a very famous paradox called Morave paradox. I don't know if you know about Moravec. He was also a CMU professor and I'm also CMU professor so have to quote him for sure. He was one of the founding figures in in AI. After 30 years of working in robotics he arrived at a very simple conclusion. Hard is easy, easy is hard. Okay. Whatever human believe to be hard is extremely easy for computers and vice versa.
And you can see it happening right in front of your eyes. You would say doing math is hard. Math Olympiad gold medal is really really hard. Climbing a stair oh super easy. Now look around you in technology where we are in terms of what has been solved and what's not being solved. So this is robotics for you. Okay. It's very it is not yet another application of AI. It is what AI was founded for in the very beginning and has made very little progress towards.
This is not yet another application where deep networks can come in and attack the field. This is has to be thought of with fundamentally first principles from the ground up. This is a very famous example of where you can have a you know a computer beat uh Gary Casper on chess in '90s but you can you still don't have a a a computer or robot that can pick up the chess pieces and arrange them on any chess board. Now, so how do we go about solving this?
Okay, so on a more positive note, uh we have a magic sauce for AI success, right? You get big data set, you train big models and magic happens. Okay, now we know this template is working well very well in many topics. Okay, now can we apply the same recipe to robotics? Well, it's not uh directly applicable because we have no data. There is no internet of robotics data and as I said earlier you can go and collect data uh manually on the robot called telly operation and as you notice teley operation is not 5 year old thing this is a 68 or 70 year old thing so the the funny part is if you go back and you look at the whole evolution of GPD3 GPD4 models go back to GPD3 uh three years ago seems like a like a lifetime ago GPD3 only began working and caught people's attention when you could train those models on trillions of tokens.
So GPD3 was already trained 30 trillion plus token and today hundreds of trillions of tokens and if you collect data by manually by tell operation it takes you about 1 minute to get one example. You can do the math if you hire all of US population it will take you more than a century to reach the same scale as of GPD3 uh which is like uh and nobody uses GP3 today. Okay. So this is very slow and expensive. So in robotics I would argue nobody has really scaled robotics yet and we are very far from talking about scale in robotics with the way other uh other areas have seen scale.
Okay so this is where uh we have been focusing on uh our scaled is about uh 3 years old but uh I have been working on the problem for more than a decade uh that's all I've done in my career uh nothing else. Uh so at scale what our thesis is is to build what we call an omniodied intelligence any robot any task one brain okay any robot it can be a humanoid it can be a quadriped it can be a robotic arm on a conveyor belt or a dextrous hand doesn't really matter this even this hypothesis is way more general than uh one would argue humans are because we control our own body and why do we have to go so general well the argument is in robotics there there is no data anyway.
So I can't pick and choose which hardware do I use data from. And we should be able to use data from any kind of hardware, any kind of task, any kind of scenario. And that's the only way to truly achieve the scale of uh of what language models achieved uh three years ago. Okay. So the goal here is this uh this whole idea of any robot any task one brain. And through this talk I'll hopefully convince you why this is the this is the way to go towards robotics.
Okay. So this is the uh rough intro but before I go into the more details let me show you just a teaser result. So in this result, every single robot uh it's from a different company, different hardware and they're all controlled by uh skilled brain whether it's humanoids going up and down stairs, any kind of scenarios. And these are not new results. They're like a couple of year old uh results in here. Robust to dust disturbances.
You can put them zero shot in new scenarios. So the idea here any every single robot in this video takes any action anywhere in the world the brain behind the scene [music] improves because it's an omni-body brain. >> Yeah, this is really hard right because like eggs are okay. So that's a teaser of what we what we work on. So I want to make this talk more scientific and more uh more informative than a company ad. So I'll talk about uh how do we scale data in robotics.
Okay, let's take a tour back as to how have people addressed this. uh I'll take a look back at my own career and my hypothesis for data has been changing over the years. Okay. When I began uh uh working scaling things the idea was you can have robots you can collect data for robots manually but it is too slow and robots should be allowed to collect data by themselves. So we had this whole idea of curiositydriven exploration.
Uh allow robots to explore in a curious way in the environment and it has a lot of good parts like no human required. Robots can go and keep playing collect more data. We called it play data at the time. It's very rich because it has force sensors and joint angles. But the difficulty is it is only in the physical world. So it's very difficult to scale. E Google was at it at the time. Uh h having hundreds of robots. Even for Google it's too expensive and too slow to scale.
The other idea is again telly operation. As I said it's an old idea but again the same issues impossible to scale because now you don't need robot you also need humans. Uh so it's even more expensive and diversity is very limited because even if you put a robot in one setup with a human you can get thousand examples but they'll all be in the same setup. So what you ideally want carry the robot to new home every day in new scenario and that's just impossibly hard to scale.
Then we had major breakthrough in learning from simulation. This was one of the first result where uh one could show deep reinforcement learning uh trained in simulation transfer to real robot. This is one of one of the award-winning paper uh at the time and now it's used in every humanoid uh every company out there. uh this was from our lab at Berkeley and CMU. But again they there are pros and cons. It's very easy to scale but diversity is hard to get because you have to engineer every scene in simulation manually.
So if you're listening to this like there is no single answer I'm coming to. Uh there is no uh no no uh no single solution. This is uh another work we did earlier. This is all before skilled learning from human videos. You watch the human do things and then robot copies this. Again high diversity. You can use videos from YouTube etc. Highly scalable but it's a very poor form of data because it's very far from robot. Okay.
So what is the solution here? The answer is there is no there is no golden path. You have to think about data in the context of different features. And in my opinion there are only three features that matter in in robotics. scalability, diversity and how close you are to your robot uh uh robot joint angles. Okay, scalability means can I quickly scale it across scenarios. Diversity means can I get diverse data because just having 100 trillion tokens is completely useless if they're all in the same environment and same scenarios.
Okay. And closeness to robot mean how far are you from the ground truth of robot own joint angles. So if you look at simulation very scalable low diverse uh uh but moderately close to robot human videos are very scalable and diverse but very far from robot data you have to learn to map the human to robot this other two teleop and the and the manipulation interfaces they are kind of in between and you can see here uh right now around the world if you look at different companies they're all focusing on one of these approaches most of them are on teleop or or yumi setups very few on everything else and but there There is no golden answer like they every each one of them have a downside but the golden light here is that they are all complemented to each other like they are not their cons are not exactly matching from each other.
So this is where uh the recipe that we have converged to over the years is to realize separate the training into two parts pre-training and post- training but that's not surprise right but how do we pre-train you want to pre-train using data which is highly scalable and diverse but maybe low quality so simulation human videos etc. So that's how you pre-train and then for post training you use data from telly operation.
Now what is teleoperation data? It's very high quality but low in amount. Okay, high quality low in amount. It exactly reminds us of the recipe in language models. You pre-train data on the internet then you post train for coding for uh uh for your own company etc. But in robotics there is one more bucket which is deployment data and deployment data is highly scalable once deployment scaled. So over time this data will take over everything else.
And what we are trying to build is what we call this data flywheel which goes which takes this data from post training time and puts back in pre-training. And now you can see why this idea of omniodied brain is extremely important because this is very unlikely that you have only one robot only one version deployed forever in every task around the world. This never happens in any area of hardware except chips because they are very hard to manufacture.
Uh and you can still see even in the chips uh inference chips are coming left and right for many companies these days. So in hardware this is never the case. you have only one version of the robot, one shape, which is why omniodied intelligence is the enabler of what we call a deployment data flywheel. Okay, so let's now look at a few results uh very quick. Now this brain is extremely general. So we can do variety of tasks very quickly.
You may have you have se you may have seen many tasks like laundry folding etc. It's very popular task in Silicon Valley for some reason. uh and uh and the and the argument here is uh when have you ever thought while folding a t-shirt that if I miss my hand by 1 cm my t-shirt fold will be a blunder you don't think like this people don't even think while folding they just hold anywhere you just do something so that the tolerance for error is extremely extremely high then why is this task hard anybody why why do people get why do people get fooled into believing this task is Let me put it that way.
Fabric, right? Why is fabric hard? Simulation. But nobody's using simulation anyway. This is all from telly operation. Why is this hard? You know why is it hard? Because it is hard for classical version of robotics. Classically in robotics, people would model the whole physics, create models by hand and then do this. It's very hard for that. But for deep learning based robotics, this is the easiest task possible because it has high tolerance for error.
So it is completely uh now so many companies focusing on this task. Now what is hard is I would say something about like let's say this task. If I ask you before seeing this video is this task doable without having hands very likely half the people will say no because it requires putting like how many of you have lost airports? Uh like and in our company there's a channel called right airpod because people keep losing their right.
Now in this case the robot does not have hand it has gripper. So here the task is very hard for the gripper. So it requires a higher level of intelligence because the arm has to go and orient itself to pick up the airpod in the right manner such that it can be inserted because the grippers can only close up and like like this. They're parallejo grippers. So you cannot turn the uh the the the airpod at the very end. Now what you are seeing here these are not real airpods.
They are fake ones from teu uh like $5 $10 each. So they don't have a magnet inside. So the robot has to really work hard to put this inside properly because there is no magnet to pull it uh easily. So these are all fake uh fake ones. This is even harder than it appears uh in the video. You can also uh like once you build the general brain behind the scene, you can also go and learn it from variety of just human videos without actually having any finetuning data at tell operation time.
So in this scenario what we did, we trained on human videos like this like egocentric videos. we are now trying to transfer it to more third person uh uh videos and if third person works you can learn from YouTube any any kind of open source data. So in this case we see the video and then we add less than 1 hour of robot data. So very quick uh transformation and then the it can transfer to humanoids. Now these ones have hands.
Hands or no hands is a big debate. People often use hands as an excuse as to why robotics is not here but that's not the case. It's always the intelligence uh behind the scene. We don't deploy hands right now because there are none available which can be deployed in factories. They all break within 100 yards. Uh pick anyone. If you have better hand uh I would love to buy 10 units ASAP to test. So it's robust to scenarios.
Now the idea is you can do many tasks by watching humans. And the reason we can work with very less data which is less than one hour of robot data is because robot imagines things in its own head and tries to multiply the learning from for many scenarios. So what you see here is a completely fake video uh you can call it robot's dream. So it's inside the robot's own model where it can imagine scenarios uh in so this is not real data.
This is all fake uh from own robot's own model. You can transfer it to you know more complex task and even lower cost hardware. So this is a egg uh uh egg making like omelette making task. Uh now here the these this whole setup costs about $4,000. So extremely cheap arms uh compared to the uh what you see out there. And if you notice here this setup is too cheap to even have any sensors. So the only sensor here is just a camera.
No force nothing else. And the robot can do task which require forces uh from vision. Now I'm not saying that's a future like you should not be doing this. I'm sure the sensors will improve but >> from the existing sensors we are way behind than where we can be from just intelligence perspective. >> Okay. So this can keep on going. This is my adviser from Berkeley. He did not believe the robots can do it. He just kept standing for like uh 1 hour and the robot kept making omelette.
Uh and and the interesting part here is this is fully end to end system. There is no state machine nothing. So the reason robot is looking egg because there is an empty plate in the front. I don't know why it's fluctuating. It's not in the there's no cut in the video. I I assure you of that. It's HDMI problem. So as soon as you remove the plate uh sorry I don't know why this is happening. Yeah. So there's no state machine when the robot starts when the robot ends.
It's all automated and it's very robust to even disturbances. You can change objects around. Uh you can add these are all unseen and unseen objects. Uh only the pan and the gas stove is same. Everything else is different and the robot can keep on going. So you can get basic robustness. So this was this is all trained with less than 10 hours of data. Uh so it's very low data to be learning this robustness and it's coming from the base model behind the scene.
I'm skipping uh this in the interest of time. So unlike traditional robotics pipelines where you have planning, mapping etc. This is an end toend brain. Now end to end is heavily used term in in many areas of AI. But when I mean end to end, I really mean end to end. It reads directly from the cameras and it applies power directly to motors. So we use nothing uh in between except just a P at the very end to go from 100 to 500 Hz.
But P is not robotics contribution. This is before it predates to like World War II or something. Okay. Now the the the interesting part is vision for us is yet another input. So nothing that crazy about it. So if you have seen, you may have seen a lot of lot of results of robots dancing, doing karate, kung fu, backflip. Go back and think how many results have you seen of robots going up and down stairs. Very few. Like even the top companies out there in humanoid, they have not shown beyond one sample stair just to check the tick tick box.
You do. And people talk about humanoids are coming, China versus US, all of this is sort of BS. If humanoids cannot climb stairs, what's the point of having legs? All right, that's the only reason why you have legs. Now, this is actually a Moex paradox add action. Left one actually is much easier for a robot. A backflip, dancing, etc. is much easier for the robot. And you may think, okay, why is this paradox exist? Think one level deeper.
When a robot is doing a back flip, it has to only know about its own body, nothing else. Right? So it's and when everything is known or fully observed that's what computers are good at because they can do they can simulate every possible uh uh setup but the right one even simply climbing on stairs requires seeing the stair at the first at the first place or what's the height what's the width you don't measure height and width exactly but you adjust to all the disturbances and that requires vision so right one requires understanding and that's why it is hard and you don't see any of this but for us it's just get another result it doesn't really matter so we And uh we put this result out one and a half year ago.
We shot this like two and a half years ago. Uh but uh like this robot can go any scenario. It is not stumbling on things. It intentionally steps over things and there is no mapping or planning. It doesn't make any 3D map of the surroundings. It's all operating from camera on the torso. And you can put this in any kind of setup, any kind of stairs. Uh I'm going faster here. These are, you know, uh fire escape stairs. It's a bit weird.
Fire escape stairs are to be used when there is fire in urgency and they are the hardest in every building. Uh so we are testing in all in all those setups but you can put this in anywhere you can disturb it and stay doesn't really matter. This is superhuman capability because robot cannot see it's being pulled. So it's a surprise factor for the robot that you're pulling the violet while climbing. And again it's the same model everywhere in all these scenarios.
So very easy although easy parkour does look fun uh uh does look fun and people often find like uh this is uh you know the robots can do all this but this is so much easier than what I showed earlier and even if you see these videos right now many of you will find this more impressive even though I'm telling you this is easy it takes half an hour to train this while the previous one it takes much longer uh and much more difficult to train so you can take this Based model and you can transfer to variety of tasks very quickly.
So here this is a very old result from one and a half year ago uh for inserting LAN cables etc. We have very advanced systems now but we are deploying these models already across variety of application. So for instance because to to build the data flywhe the hardest part is the deployment itself. There are so many other issues beside intelligence and hardware that come up when you deploy robotics. It is not the same.
This is not same same as deploying chat GP over an app because you have to work around many of the and I think Skyio gave a talk before this and they can talk to you about all day about the hardness of deployment. So we have to start deployment now so that we can have this data flywheel in a foreseeable future and I'll give you a few examples of deployment we are doing. So one of the examples with Nvidia like Nvidia is opening their first factory in Houston. uh and their GPUs what you use right now they're all built outside US in in Taiwan mainly and it's all done manually over there and uh humos are really efficient and really uh really good at making these things but sustaining cost and scalability and throughput here in US it's very hard to maintain without this labor force so here we are automating uh this GPU assembly for them this was we did a live demo in Nvidia GTC this is deployed in factory already uh last week so it's already live but the interesting part I want to highlight this is a very traditional factory task but if you look at the the whole setup this is extremely randomized extremely noisy and you go to any factory traditional setup they're extremely clean so here the robot the same brain can not only do very precise task it can be very robust to any kind of disturbance which means now these robots do not need to be in a cage uh and they can be uh just deployed as is on on these factory lines without any change to any factory line alongside humans and wherever there is humans there is mess.
Okay, you can put the same brain to delivery applications. So here the robot delivers from the from the truck to the to the door. It has to find where the front door is. So it's a common sense problem. It's not mobility problem. And we already have been delivering uh packages to with working behind the scene with many partners to delivering packages to front door of the houses uh in these areas. same setup works across warehouses and all this.
Um now just to close the talk uh close the topic I mentioned this is an omnibody brain and I hope it's very clear why it's extremely essential to have an omnibody brain to build data flywheel but there's one extra benefit and the benefit is your robots may break over time and uh or the safety uh scenarios and safety is a byproduct of this omnibody brain because if you have a we can put the I'm just skipping very fast here so you can see the online this is all online uh but we can put the same brain across each of these systems each of these robots and in this case we did not even train on these robots.
This is completely zero shot transfer to all these humanoids and all these things and even if the robot breaks like here the leg breaks it can recover within milliseconds because when your leg breaks your three-legged robot is a new robot. So it doesn't really matter what's the shape of the robot is anything changing here we disabled the legs of the robot it learns to learn it learns to walk on two legs in just three trials.
So what you are seeing on screen is all the training for the robot that they are that that is happening in these systems. So it's it adapts in few milliseconds to uh uh 30 seconds or so. And like one example here it's it's going on wheels. You jam the wheel it starts walking. So this is a another byproduct of omnibody brain. Safety takes a very different meaning uh with these kind of models. when the robot can fly like or sorry when the robots can walk or or operate with only half the body.
I removed some of the gory scenes from this but we also cut the robot in half and it can still work uh with both the two halves separately but for that you can go to YouTube uh and that's all I have. Thank you.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.