Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 11:58
4.6x that video's typical replay level
issues. Uh I also invented OS certification. I just close the tracker whenever I want, so I have my life back. So, does this work? Yes, sort of. >> [laughter] >> Which leads me to act three, slow the down. Everything's broken.
Said at 11:52
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
The graph counts replays. It does not show where viewers stopped watching.
Words
3,438
Runtime
19:01
Speaking pace
181wpm
Reading time
14min
181 words per minute, the same as the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
All right. Hello everyone. Um, my name is Kevin. I'm going to be talking about anti-gravity. So, are there any World Cup fans out there? Woo! Imagine you are coaching Argentina and you're in the 89th minute and you have Messi on your team. What play are you running? It's called Give Messi the ball and get the heck out of the way. LLMs aren't just role players anymore. They can be your star player if you build the right product around them. And to let your star player cook, you
91 words, the words spoken in the first 30 seconds at 181 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 208 |
| Average words per sentence | 16.5 |
| Longest sentence | 56 words |
| Questions asked | 9 |
| Sentences containing a number | 25 |
Most used terms
Filler phrases
73 in total: um 17 · actually 15 · like 11 · sort of 9 · uh 9 · basically 4 · kind of 3 · you know 3 · right? 2.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
All right. Hello everyone. Um, my name is Kevin. I'm going to be talking about anti-gravity. So, are there any World Cup fans out there? Woo! Imagine you are coaching Argentina and you're in the 89th minute and you have Messi on your team. What play are you running? It's called Give Messi the ball and get the heck out of the way. LLMs aren't just role players anymore. They can be your star player if you build the right product around them.
And to let your star player cook, you have to get out of the model's way. We might want to get the slide. Are the slides up? Oh, they are. Great. Um, so Anti-gravity is Google's agent coding product for technical and non-technical users. Uh we launched back in November of 2025 and have been accelerating devs both within Google and externally ever since. My name is Kevin How and I lead part of the engineering team on anti-gravity.
So let's talk a little bit more about what anti-gravity is. We have and always will be unapologetically agent first. So, we debuted the anti-gravity IDE last year with a brand new agent manager concept and it was a platform to manage and orchestrate many agents. Since then, we've actually extracted our agent and launched our own anti-gravity CLI. And last month at Google IO, we had the pleasure of launching anti-gravity 2.0.
In the theme of getting the model out of the way, we actually decoupled the IDE from the agent manager. So, now you have two separate applications. Um, and now you can use the agent manager in a standalone app. And since pictures are worth a thousand words, here's a screenshot of anti-gravity 2.0 in action. As you can see, not only is it your own dedicated mission control for your agents and projects, you have sub agents, you have all the new models, you have work trees, scheduled tasks, voice mode.
There are so many things to unpack with the product. But I don't want to spend today telling you about the product. I want to tell you a little bit more about the behind the scenes, some of the principles that went into it and notably some of the things that led to its roadmap. So, as some of you for the longtime AIG fans, uh this is actually my fifth time speaking at AIG. Um and I've been building developers tools since 2022.
And the one thing that has stood above all other lessons that I've tal talked about is the idea of scaling with intelligence. This means that as the model gets better, so should your product. And the frontier edge of whatever model you are serving should be apparent inside of your users's product experience. So let's get into more concrete examples of what this means. So for those of you that follow me on X or hear me just yap generally for the last four years, you'll know that I've been working on a number of these sort of transformations year over year over year.
In 2022, I was working on autocomplete and chat sidebars. This was based on embeddings, rules, files, a syntax tree parsing, basically everything inside of that app is deterministic because that's all that the model could really handle. And in 2024, when agents came onto the scene, it completely changed how developers were going to do work. With it came new primitives like MCPs, custom tools, and permission systems. And with 2025, we introduced anti-gravity's agent manager with many other products following suit in that similar form factor with users managing many agents at once in parallel.
And this led to things like skills, hooks, artifacts, and a couple other primitives. Um, and that sort of defined the 2025 era. So, let's talk a little bit about 2026 and what those primitives might be. Before we answer this question, I want to take you back to some of these battle scars that are a little bit closer to home. Scaling with intelligence really is not easy. It's really hard to take away something that users love and are familiar with to lead them down potentially, and that's a big keyword, a better path.
We aren't right 100% of the time, but there are two that jump to mind when I was putting together the slides for this talk. The first one is giving AI a terminal. We all remember fears about son of Anton deleting your entire codebase and doing catastrophic things to both, you know, your startup, your company, etc., etc. But as models got better and people invested in primitives such as permission systems, users ended up building faster, they ended up shipping more, and they did so safely.
So, we were able to overcome this. And as models got smarter, they were able to make better decisions about what they should and should not run in your terminal. The second instance is um this tweet which is very representative of sort of the yelling that I got uh when we removed chat from windsurf. So a lot of users were yelling at our team because we took away something that was very dear to them the chat sidebar and replaced it with only an agent.
Now at the time this is something that was familiar and rather difficult to swallow. But when we look back, models have advanced. Multistep research, agentic research and execution became the new paradigm. And here we are today using and loving all these agentic products. And so now I bring you to today's battle. What is going on today? So we decoupled the agent manager from the IDE. And with anti-gravity 2.0, we split them into separate applications.
We believe that the IDE is to the agent manager what the debugger was to the IDE. You don't always need a debugger, but it definitely is helpful to have it if you need to go a layer beneath and go one step deeper into that abstraction stack. And our prediction is that this idea of agent orchestration, you can call it agent teams, you can call it swarms, you could call it software factories, is the future. And we're willing to bet on that future.
So here are the primitives for what we're calling the agent teams 2026 era. These are things like sub aents, generative UI, and sidecars. And we'll talk more concretely about what those things are and some examples of how they manifest inside of the product. But it's really important to first understand the why. What brought about these changes and what model changes, what model properties actually led to the development of these new things?
And as a product team, do you force the new era of primitives or is it something that comes to you by using the model and experiencing the model? The answer is kind of both, right? And the privilege of being inside of Google DeepMind is that we do have that relationship between the product and the model. So you remember the crux of anti-gravity 1.0 is to manage agents in parallel, to put the human in the driver's seat.
And if you remember my last talk, I talked a lot more about this research product flywheel. And now, as promised, because of the anti-gravity product, Gemini has now learned a thing or two about how to manage a team of agents. There's still a lot of headroom to make multi- aent systems better, more collaborative, better at deconstructing tasks into smaller tasks. But we've got a really good head start with Gemini. and all the basics have been imbued to the model so that we can build a product like anti-gravity 2.0.
Gemini 3.5 FL yeah Gemini 3.5 flash was launched back in April and this brought to market a lot of those capabilities that we had been working on in the background with anti-gravity and flash now isn't just good at executing tasks. It's actually really good at leading teams. It's faster and cheaper, pushing the paro curve of what is intelligent versus the speed and the cost at which you run those things. And putting this all together, we were really excited to announce agent teams in public preview inside of anti-gravity.
All you have to do is simply type the slash command/teamwork and you'll see a new mode where you can enter and unleash a swarm of agents onto the task at hand. So we'll talk a little bit about how this works. You as a user will specify your task. The more specific you are, the better. Though the nature of these agentic communication styles is that if it needs something more, it can actually ask you for more until everything is basically clear.
You'll work with that lead agent and it will manage a team of arbitrary size to get that work done. And what I like to say, it's kind of like the Avengers, right? It'll take a bunch of specialized roles. It may front-end engineers, backend engineers, infrastructure specialists, QA, design, the list goes on and on and on. And there are infinite possibilities for what each of those sub aents could take on. Each sub aent is dynamically generated and can operate independently.
Um, and it can even actually select a different model from what the main agent is using. And this is done so by that main agent. Again, we are scaling with intelligence. And one of the coolest aspects of this is that it can use generative UI. With a model that is as fast as Flash, things can happen nearly instantaneously. If you ask, hey, what is the status of my task? show me a canban of what's going on. Or maybe, you know, you prefer something a little bit more like uh the the Chrome debugger tool.
It can show you a timeline like that. And all these things are generated on the fly because it's able to generate UI on demand. So, some of the projects that the system has implemented, um we've built a photo editor. You can actually edit raw photos directly inside of your browser. Um we've also built a messaging app that might look a little bit familiar to those in the room. Um, and each of these took hundreds of sub aents uh, and took almost half a day to run.
But to really put it through its paces, one of the hero runs that we did was actually building an entire OS kernel. This is something that, uh, we got to show off at Google IO, but we built a complete OS kernel from scratch and actually played Doom on it. And my colleague Verun was able to demo this at Google IO. We were super proud of this particular milestone because it really demonstrated that if you throw more intelligence, you throw more sub aents um at this sort of problem, a model like Gemini 3.5 Flash could do this in a way that was not only very very powerful, but also scalable and you know, mildly affordable.
Obviously, we're not going to spend thousands and thousands of dollars to build an OS kernel every day, though it is possible. And some of the stats out of this, it took 93 sub aents over the course of 12 hours, made 15,000 requests, two billion tokens, and it was under $1,000, which was one of the really cool aspects of this project. And so, as you can see with this particular example, sub agent primitives are one of the defining parts about building a 2026 era of agent teams.
So, agent teams are just that first example. And I want to show you another example that our team uses internally that sort of demonstrates some of these new primitives. Um the second one is about automating research tasks. So we work inside of Gemini. We help sort of make Gemini better at coding related tasks, agentic related tasks. And this is where the real magic starts happening with the product. We have an internal version of anti-gravity that researchers, engineers, nontechnical folks can use.
And when they understand the primitives that anti-gravity offers, it becomes a very very powerful way to automate your own workflows. So we'll take the example of sideby-side eval. So this is a very common workflow not only at DeepMind but just generally in the industry. You essentially will take multiple rollouts 1 2 3 4 etc. Um and you want to compare them. So you'll take a set of tasks, you'll do some rollouts, you'll get some results and they'll essentially be in two different tables.
Now you'll look at the control, you'll look at the experiment, and then you'll have to figure out not only what the difference was, but perhaps what are the reasons for those differences and how can we actually iterate from there and make a better version of for the next experiment. Now, traditionally, this was a lot of Jupyter notebook elbow grease essentially. But when you start working with the new primitives in 2026, you end up with a lot cleaner of a workflow.
So researchers were able to automate 90% of this workflow by simply asking the agent about the eval in question using natural language. Then the agent that is now primed with skills and an understanding of Google's massive monor repo codebase is able to crunch the numbers and get back to you with a delta. Now what's really cool here is instead of just taking that delta then handing it back to the user, it went the extra step.
It spun up for a research agent specialist that proposes a hundred different hypotheses over why those deltas might occur. And then it uses sub aents to then spit up one sub aent for each hypothesis and basically drills into that particular case in parallel mapping back to a single response and then telling the researcher, hey here are some areas that I found. Now uh here's a report that you can review. And what's really cool is that it doesn't stop at just the report.
It actually puts together a generative UI for you to look through, interact, select dropdowns, filter, segment, slice, and actually interact richly with that data. And internally, we care a lot about this sort of workflow. Improving the model, improving the product, and understanding the ways that users find success and failure internally at Google. So, what used to be a very manual process now takes minutes. So, what used to be handineering, you'd have to build your own async pool of agents, you'd have to set up your judges, uh you'd have to tape together data pipelines, all of this now starts becoming grounded in these new primitives that we've established earlier in the slideshow.
You have a sub aent graph that is completely dynamic. The generative UI comes in at the end to richly convey the findings in a way that the user best understands or maybe caters to their learning style. And all of these things can be regenerated and redone on the fly. All the user had to do was load up a skills file and ask away. So with teamwork and this eval example, we start arriving at these 2026 primitives that I keep talking about.
And these model characteristics really change the way that we have to think about the product and how we have to develop the product. So the three examples that we've talked about, first we have the dynamic sub agent. And to provide a little bit more color here, basically no two sub aents are the same. The main agent is the one that is orchestrating this entirely on its own. It's configuring and prompting and seeding these sub aents on the fly.
They can operate in parallel. They can operate in different types of secure environments be it a sandbox be it a remote execution system and they can all take on infinitely an infinite number of specialized roles. So the scaling story here is quite obvious and from the last two examples you can probably tell as the model gets smarter your team will become more specialized it will become more collaborative and ultimately that means it'll be capable of getting more complex work done for you.
And now the second is this new concept and we've alluded to it slightly in the past, but it's called sidecars. This is a new plug-in protocol that we're bringing to anti-gravity. A sidecar process is essentially it is a sidecar process. The naming sort of reflects what's going on under the hood, but it's a longived utility and it's responsible for listening. It allows the model to listen to the outside world and set up its own triggers for things that might happen.
For example, this could be SMS messages. This could be web hooks, cron jobs, hooking it up to GitHub PRs, the the list goes on and on. But this is a generic plug-in primitive. Anti-gravity already uses sidecars for things that are timebased. This is where the the scheduled task cron concept comes from. Um, but under the hood, this is all this new sidecar primitive. So, we'll be releasing the spec for this so that you all can build on top of this new primitive um later this summer.
But there are some really, really cool ways that people internally have been using this sort of concept. And the third and final primitive is generative UI. So we hypothesize that human written specialized UIs are kind of dead. Gemini Flash on anti-gravity clocks in at almost 900 tokens a second. This is 10x faster than a lot of the other frontier model experiences. And in a matter of seconds, you're able to go from whatever you thinking inside of your head into a prompt into a use case that is designed and embedded inside of your conversation view perfectly.
And rather than rely on templates or even HTML files, anti-gravity can render your generated UI in line. So you can do things like this and play Doom. But this also extends to things like bar charts, graphs, um, tables, anything that you would want to interact with and maybe inspect a little bit further than just a markdown file or just a conversation. And generative UI in many ways reminds me of the quote that the late Steve Jobs said when unveiling the iPhone.
He justifies the removal of the keyboard and says they all have these keyboards. They are there whether you need them or not. And they all have these control buttons that are fixed in plastic and are the same for every application. In an analogous way, we built our product to dynamically scale with the needs of the agent. We skipped the heavy infrastructure and mechanical UIs in favor of sidecars and generative UI. And that creates a product experience that is not fixed in plastic.
So, sub aents, sidecar triggers, and generative UI are the latest primitives that are powering anti-gravity. We've tried our best to stay out of the way and let the model cook. And if you're building a product around an agent, you should consider what are the primitives that are in my product and how might they scale with the model's intelligence. We all are familiar with shipping features is now quite easy with all of these new tools.
And it's about deciding what features to actually add um so that the model so that your product can scale with the next release of the next model which will inevitably be faster, better, and cheaper. And so with the right primitives, you as a builder or you as a product owner, you might be surprised at what the models can do. And in classic fashion, I'm going to keep using the slide until we've actually conquered the TPU crunch.
So you can find me on Twitter. Uh you can DM me for feedback. We're always looking for new ideas on how to build the latest and greatest. Thank you for watching. Thank you Swix and Ben for having me. It's always a joy to be here. And I'll be at the anti-gravity booth if you want to talk further. If you want to get to know the product a bit more, some team members will be there. So, thank you so much for your time. Excited to meet you all.
Yeah. Heat.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.