Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
Most replayed moment at 2:02
2.5x that video's typical replay level
are re-bumbling around the work itself. And the important question here becomes a lot less about what is your title and more what part of the system can you own? Now, I like this taxonomy quite a lot.
Said at 1:56
The graph counts replays. It does not show where viewers stopped watching.
Words
2,695
Runtime
19:43
Speaking pace
137wpm
Reading time
11min
137 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
[music] All right. Uh, I understand that I'm standing between you and the launch, so I'll try to be quick. Uh, my name is IU. I'm a professor at Ohio State, the Ohio State and uh I also have another job which is a CEO at a company called Neocognition and we focus on agents and container learning. So today's talk um it won't be too technical but I
69 words, the words spoken in the first 30 seconds at 137 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 103 |
| Average words per sentence | 26.2 |
| Longest sentence | 153 words |
| Questions asked | 22 |
| Sentences containing a number | 4 |
Most used terms
Filler phrases
219 in total: uh 101 · like 59 · um 26 · right? 16 · actually 10 · kind of 4 · you know 3.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
[music] All right. Uh, I understand that I'm standing between you and the launch, so I'll try to be quick. Uh, my name is IU. I'm a professor at Ohio State, the Ohio State and uh I also have another job which is a CEO at a company called Neocognition and we focus on agents and container learning. So today's talk um it won't be too technical but I will it will be mainly a a conceptual one but I think it's a very important uh conceptual distinction that I will try to make between what is intelligence and what is expertise and through this I will try to answer some of the very bothering questions for me that um like why we are so successful at the coding agent but they're so terrible at anything else, right?
Why uh the current agents are so token inefficient uh that to the degree that every company right now is like coming out and try to curb their uh their token maxing efforts in the company. Um so hopefully this will provide some food for thought before launch. Right. First a bit of a history. Um so AI agents are not a new thing, right? it's uh we have been trying to develop agents throughout the whole history of AI. Um but the problem is that in the early stages uh let's say um in the n in the 1960s to uh 80s when we developed these expert systems or logical agents or like in the uh 20110s when we developed these deep RL based uh neural agents we were only able to capture some very limited facets of human intelligence right whether it's like logical reasoning or it's like uh perception in single modalities to decision.
Um only recently with the mult model LRMS and the language agent built on top of them for the first time we have a neural model that is able to encode multiensory inputs into a uh unified neural representation that is also conducive to symbolic reasoning and communication. Right? So that was a trait unique to humans. Now uh AI agents finally have the same thing. So that drastically improved their expressiveness, their reasoning ability and adaptivity.
So that's why I think we have really entered a new evolutionary stage of machine intelligence. And uh it didn't take long for these language agents to find their first mass markets which is coding. And the best way to illustrate this is probably through the uh revenue graph of anthropics right in just under two years their revenue has grown 400 times uh to uh 40 billion I think the newest number is maybe 60 billion uh annualizer runway and it's largely driven by coding and coding related productivity uh capabilities but if we think about it right coding is the really the ideal market for this language anguage agents because code is already a language native world.
Everything is already represented symbolically and like uh recorded uh in a very structured way and you get your rewards, you get your uh like tests all in place in symbolic ways. So then what happens when we leave the privileged world of code? Well, not so well. Um we are running into a lot of challenges deploying these agents in enterprise settings and the uh also in personal settings like the uh like open clause uh constantly make this uh like uh quite brittle and silly uh errors and then to the extent that Andrew Capathy uh said that it's not going to be the year of agents it's going to be the decade of agents because they cannot do computer use they don't have continuous learning um I don't know how much uh thought has changed uh since like uh last time because of the coding agent everything but I think the difficulties with computer use with container learning still largely the same right now so how can something be so small but also so brittle at the same time here's my thesis around it I think we are actually witnessing a modern version of the Maravx paradox, right?
So the par morave paradox essentially says that for AI uh hard things are easy, easy things are hard. So the modern version here is that we are very good at this symbolic reasoning task like coding and math which were considered crown jewel of uh of intelligence uh earlier. But then we still struggle with this everyday digital work because they really require quite different set of cognitive competencies to excel as them and more specifically I think modern society is really not just one unified world it's millions of these micro worlds like every domain every every profession is different every company is different even if you're using the same software every company company configure it differently so is extremely idiosyncratic especially in the digital world.
It has its unique local physics like different structures, constraints, affordances and dynamics that you have to learn. It's just like too hoggenous and dynamic for any uh monolithic model to try to compress it into one static representation. So agents must continually learn on the job to acquire what I call specialized expertise for each specific micro world. The second part of the talk I will try to establish the differences between intelligence and expertise.
Here are the working definitions. For intelligence, it's the capacity to reason through unfamiliar problems from available context. Right? This is what the frontier models are increasingly good at. Um, you give it the provin statements, the context, the tools, and they can reason through this even if it's seeing them for the first time. Uh, and they can do a great job. Every episode is more or less like independent from each other here.
But expertise is different. Expertise is really accumulated and situated competence. It's the ability to act reliably, efficiently and with judgment to achieve reproducibly super real uh performance in a particular domain. Right? So this is in stark contrast with intelligence. And to uh show what does expertise actually contain. I think uh the key idea from cognitive science is that experts don't just know more facts, we actually see the world differently, right?
So um the uh expertise allows you to do different pattern recognition. So you see through the surface patterns like if you're looking an expert is looking at like a gigantic bug reports they can immediately locate like the most plausible places where things could go wrong. Um and you think about the problem with like a very deep structure, right? When you are scheduling a meeting, you know that it's not just like finding the shared slots on everyone's calendar is actually a constraint optimization problem over everyone's authority, the priorities, the urgency and everything.
Um and we don't expert don't just operate with a set of rules, a set of facts. We know that every single thing is conditional, right? Every rule has like the preconditions where it applies. But we also know when we can bend the reality, we can bend the rules when exceptions happen, right? And finally that also give us judgment and taste is importantly what's high quality and uh very importantly when to stop when is good enough.
Um so all this together I think experts effectively has have built a world model of their environments right that it's a generalized notion of world model that captures how that micro world works and that becomes the basis for all of our perception uh reasoning decision- making uh and judgment. So intelligence and expertise are really quite different across many dimensions. Uh but some of the import uh interesting ones here are like intelligence is about hey when we have the context uh how to solve the problems through the context but expertise actually will bring you the right context right given any problem we know what context bring into are important for this problem and bring it in to solve the problem and because of that uh intelligence tend to expand our search like every problem solving is a search problem.
So, intelligence tend to brute force it, try to uh try to spin up like 100 different uh like parallel ways to to try to solve the problem. Well, expertise will actually try to compress the search space because expertise has constructed this has learned this essential shortcuts for the problem space so that whenever you have a problem, you know the most plausible ways to solve it. Um and then I also think the final part here is that I think continuous learning is the important bridge from intelligence to expertise.
But first let me try to define continuous learning because it's such a confusing term um and and Jack just uh gives some definition earlier uh with like 10 different names. Um but here's my the definition I work with. I think continue learning is adaptive compression of the experience into reusable structures for future behavior. So all of these four elements here are very important. For experience, we need to uh answer the question like what kind of experience we're talking about is more like episodes of experience or it's like these semantic facts or procedures or feedback from human or environments.
And how do we compress that? Uh so in we embed them into vectors or we index them into some symbolic structure. Uh we uh distill them into model parameters or do some kind of reinforcement learning. And it's not just like one time compression. It needs to be adaptive compression like what you have learned what you have compressed so far should largely uh influence how you compress further. And what kind of structure we're looking at?
It's just like parameters like adapters of your language models or it's vectors, graphs, skills or even word models. And then how do you use these reusable structures? It's like you use it just to recall these facts or use it for prediction of like future states. is use for uh for better planning for or even for the control like actuation layer of the agent or as a value function for potential states right so it's because of this uh the continent learning problem is so rich like it has this four different aspects and if different aspects can be instantiated in different ways that makes this field so confusing but hopefully this is a definition that uh encompasses uh most of the versions of continuous learning.
Then I think that uh this is maybe the most important figure in this talk. Um if we put raw intelligence as the x-axis and uh expertise as the yaxis I think we'll find that they are largely orthogonal to each other. If you don't have continuous learning uh all you do is scaling your model to to get better like raw intelligence then what we will get is what I call the world's smartest novice right super smart it can try to uh try to attack at any problem uh provide given to it but it doesn't accumulate expertise so it end up just like brute forcing its way at every problem then if you have continued learning uh like Different continuous learning algorithms will essentially set the slope of your learning uh curve here, right?
If you have a sloppy CL algorithm, maybe some kind of simple incontext learning then uh with like uh increasing intelligence then your expertise will increase like a little bit but you have a really strong uh contin um increase like rapidly. Of course, this is assuming like a given time horizon and experience horizon. And then among all of these potential futures that good continued learning will bring us. I think this is pro probably the the one I like the most or I think it is the most interesting which I call the unbounded expertise from bounded intelligence.
Right? What if we can uh come up with a computer learning algorithm such that um given up once the raw intelligence has crossed a certain threshold we don't need stronger intelligence anymore like learning will bring us like unbounded expertise once we have like a reasonable level of intelligence right then we can call this the escape intelligence and if this is indeed true then it will have a a lot of imp implications for the whole ecosystem, right?
Do we need to continually train to train these larger and larger models or like this models like uh methods uh maybe they're already good enough? What we're missing is just like better continuous learning algorithms. So to be a little bit more concrete I think uh to provide more food for thoughts uh here are some open questions I think uh in this space the overarching question is like given any domain or environments right how can an agent continually learn to specialize and reach expert level competency but to do that you need to answer many other questions like how do you even measure uh define and measure expertise and this is probably environment specific and how to handle the trade-off between reliability and plasticity, right?
Um we want these agents to be both reliable and plastic, but they are inherently conflicting with each other, right? Reliable systems or stable systems, they resist the change, but the plastic systems likes change. So how do we reconcile that? Um but fortunately we do have a living existence proof which is us ourselves humans uh that we are incredibly plastic but also manage to be dependable most of the time. Um then uh from a technical perspective like when we talk about learning largely there are like two forms of learning parametric or nonparametric.
So how and my uh belief here is that both are really needed for uh this type of machine learning to to actually work. But how do we synergize the two and finally even though we are focusing on specialization I think there is a great potential for specialization to actually gener to lead to like better generalization. You know we are we have exhausted the public data for training LLMs but the next stage of training the next internet scale data opportunity is actually in all of these different uh like private worlds.
If we can make this specialized agent work they can learn in sichu and uh channel back the learning to the general model I think that may be the next uh internet scale uh data opportunity. Okay. So finally a call to action um think let's start scaling expertise this will be a new dimension for us to scale because intelligence is already becoming abundance the frontier models they are probably smarter than average humans um but expertise is still scarce and we want to build a world where expertise becomes abundant where everyone can gets expert uh support because in an ideal world everyone can can have their personal health care, personal financial advisor and uh personal tutors and so on so forth and then every company can build their their own learning loop, right?
They can, as Satia said two weeks ago, like we want to enable this human AI learning loop uh at each company that turns into institutional memory and for every company to build their uh own modes and to uh to still be in charge of their means of production. And finally um I think with abundance of expertise uh we will actually see more types of work become possible because there uh right now there are still a lot of opportunities that uh that are locked up because the friction is just so high to make them uh econ economically viable.
But with abundance of expertise, I think that we will be able to lower the friction and uh make many of the new type of work across the threshold of worth doing. So this is the future we're building uh towards new cognition and happy uh to share this with you and uh thanks for the attention. [applause]
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.