Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
Most replayed moment at 13:49
2.3x that video's typical replay level
Fable uh and it runs into an unknown, ask it to log it, right? So that um you uh you can see where the deviations happened and then you can sort of figure out why as well, you know? It will usually give you some context about what happened.
Said at 13:43
The graph counts replays. It does not show where viewers stopped watching.
Words
3,244
Runtime
21:14
Speaking pace
153wpm
Reading time
14min
153 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
[music] >> Good morning. I know a lot of you in this room. It's great to see you. Welcome to the voice track at AI Engineer World's Fair. For those of you who don't know me, my name is Quinn La Holman Cramer. I work at a company called Daily. We make developer infrastructure for real-time audio, video, and AI. And we're the team behind Pipe Cat, which is the most widely used framework for building voice
77 words, the words spoken in the first 30 seconds at 153 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 210 |
| Average words per sentence | 15.4 |
| Longest sentence | 54 words |
| Questions asked | 7 |
| Sentences containing a number | 42 |
Most used terms
Filler phrases
42 in total: like 21 · actually 5 · kind of 5 · uh 5 · sort of 2 · literally 1 · right? 1 · um 1 · you know 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
[music] >> Good morning. I know a lot of you in this room. It's great to see you. Welcome to the voice track at AI Engineer World's Fair. For those of you who don't know me, my name is Quinn La Holman Cramer. I work at a company called Daily. We make developer infrastructure for real-time audio, video, and AI. And we're the team behind Pipe Cat, which is the most widely used framework for building voice agents today.
Pipe Cat is open source and vendor neutral. It's used by companies like AWS and Nvidia and Anthropic and thousands of startups and scale-ups and enterprises. And today I'm going to talk about what kind of agents we're building today, including voice agents, but not just voice agents, and what I'm interested in building next. And I'm going to try to put all this in the context of the roughly 80-year history of digital computing so far.
So, we've got a lot to cover. We're going to go fast. But we're going to start in 1945 with an essay called As We May Think, written by an engineer, an academic, a civil servant named Vannevar Bush. Bush deeply understood technologies ranging from analog computers to photography to radio to radar. As We May Think is a extraordinary piece of writing. The essay predicts the development of, among other things, document display on a screen and document scanning and OCR and speech-to-text and text-to-speech and programming languages and hypertext and search engines and data networks.
Something like the GoPro camera, something weirdly like the Amazon Kindle store, and voice interfaces and brain computer interfaces. And I've been thinking a lot about As We May Think lately because Bush wrote this essay right at the very beginning of the computing age. And I think it feels to most of us like we're working right at the beginning of a new age, the intelligence age. So, what will we build? Well, at the moment we're building agents and we're having a lot of fun doing it.
And a lot of the AI engineering work we're all talking about this week is focused on building a full coherent software stack for AI agents. Here is Satya Nadella talking a couple weeks ago on a crossover episode of the No Priors and Latent Space Pod about the challenges of building agents in 2026. >> That's sort of >> That's right. So, so in some sense you kind of want to harness to define the models, the the data, uh and the tools.
And so that you have a loop across those three. And so what we are trying to first of all make sure is each of our products that we build, right? Whether it's GitHub Copilot or the security copilot the stuff we showed with M dash or even the discovery for science, it doesn't matter. All of them are multimodal harnesses um with tools access so that you can do this progressive uh disclosure of tools even so that they're token efficient.
Uh and then you're feeding it with very rich context. >> So, if you were here last year at AI Engineer World's Fair, you could draw a through line from the things we were talking about last year to loops and tool calls and context engineering and the stuff we're focused on this year to some emerging ideas. Uh you can hear that in Nadella's clip just there. I think of this is kind of agents plus plus, like multi-model harnesses and software copilot embedded in every single piece of software, organization level harnesses.
So, how do we go from agents to agents plus plus to the next thing beyond agents. Well, the last time we had this kind of massive change in how we write software and what we write software for and to do was the early days of the World Wide Web. And I was around for the early days of the World Wide Web. I was a baby programmer in 1995 and the thing we talked about all the time in 1995, the way we talk about agents today, is web pages.
I spent a lot of time writing HTML by hand and building web server software in C and indexing and search software in C and authoring tooling and management infrastructure for web pages in Pearl. I was as excited about HTML in 1995 as I am about agents today. And the web page is still with us and it's still important and useful. But today we talk a lot more about web applications and native mobile applications than we talk about web pages.
So, just like we went from web pages to full-blown web and native mobile, clearly we're going to chart a path to a new fully AI native software that comes after agents and agents plus plus. So, let's keep going back in time to in order to think about this future. Here's a timeline Vannevar Bush lays out and as we may think, he talks about the abacus, which was both an immensely useful device for doing practical everyday mathematical calculations and also an incredibly important theoretical tool that led to ideas like numeric place value and the concept of zero.
And Bush talks about the massive jump from the abacus to the state-of-the-art electromechanical keyboard calculating machines that he had in 1945. And then he posits that we're about or he is about to witness and help create an equally large leap to what he calls the arithmetical machine. And then he goes a step even further than that and he invents or designs in that essay a device he calls the memex. And we have a little bit of an advantage over over Bush in 1945.
We've seen 80 years of computing play out. So, we can modify his timeline a little bit. We can go from the abacus to the stored program computer to 40 years later the personal computer and 40 years after that this AI agents era that we're all collectively helping to invent and create and bring into being. So, the question for me is what did we build to go from those very first digital computers in the 1940s to the personal computer in the 1980s?
Well, in the 1950s the big job was to figure out more effective ways of transmitting human intent to these new computing machines. We built the first programming languages. We wrote the first compilers. And the the theoretical underpinnings here were figuring out how to combine the elegance of mathematical formalisms with something a little bit more like natural language. And then building on that in the 1960s the challenge was to make these machines interactive.
Make these machines capable of a two-way dialogue with humans. The '60s also saw the birth of graphical programming with systems like Ivan Sutherland's Sketchpad. And the '60s were an amazing era for science fiction. Even though almost nobody had access to a computer the computer became a big part of the popular imagination. The idea of a computer really resonated with people and ideas matter. For example, here is the idea of the computer in Star Trek. >> Put her on record. >> Recording. >> Come.
Captain's log supplemental. Engineering officer Scott informs warp engines Can be made operational and re-energized. >> Computed and recorded, dear. >> Computer, you will not address me in that manner. Computer. >> Computed, dear. >> I love the background sound of punch cards going through a punch card reader. So, like you know the computer is working even though it's talking to you about what it's actually computing.
There were of course a bunch of other talking computers in in science fiction of the '60s and the next year after this, the Kubrick movie that was a interpretation of Arthur C. Clarke's 2001: A Space Odyssey had the HAL 9000 computer. This is a much, much more dystopian view of a talking computer than the Star Trek computers. And by the 1970s, computers had become powerful enough that the next big job was designing abstractions that could scale to much larger amounts of data and much more powerful computing substrates.
We got relational databases, which introduced new theoretical underpinnings for data manipulation. And we got declarative languages, which leveraged those new theoretical insights. And programming languages in general continued to evolve in what to me at least are really amazing ways. We got Smalltalk and object-oriented programming in the '70s. And all of this set the stage for the personal computer in the 1980s. The Macintosh shipped in 1984.
Windows 1.0 shipped in 1985. And Microsoft's mission statement was a computer on every desk and in every home. And incredibly, Microsoft delivered on that mission statement. And we got a computer on every desk and in every home because these new personal computers delivered real, amazing, tangible benefits. Take VisiCalc, for example, which was the first spreadsheet program. A truly new abstraction for doing computation, numerical computing, two-dimensional, interactive, so durable and so useful that probably most of us in this room use a direct descendant of VisiCalc regularly, Google Google Sheets or Microsoft Excel or whatever.
Or put another way, this was a spreadsheet in 1957. And this was a spreadsheet in 1985. And I think a lot about VisiCalc these days too because I think VisiCalc is an example of how transformative new technologies can be in the way of delivering a capability that used to require a lot of specialized people and specialized knowledge and making it generally accessible. And I think VisiCalc is a potentially a counter-argument to the argument or the fear or the concern that AI is going to lead to mass unemployment.
Because VisiCalc didn't put accountants out of business. Instead, it made much much much much more accounting-like work possible. And it made new categories of work possible that we couldn't even really conceive of when a spreadsheet or doing a screen's worth of calculations as we think about it today took a roomful of people. So if we were here in the Moscone Center in 1985 and these two interfaces, the Macintosh System 2 and Windows 1.0 were state of the art, what would we have said the world would look like in 10 or 20 or 30 or 40 years?
Well, we actually have a really great example of a prediction from that time. Like as we may think another famous document in the history of human computer interaction, concept video from Apple made in 1987 called Knowledge Navigator. This is very much worth tracking down online and watching all of if you haven't seen it. I'm just going to play about 20 seconds from the middle. >> You have three messages. Your graduate research team in Guatemala, just checking in.
Robert Jordan, a second semester junior, requesting a second extension on his term paper and your mother reminding you about your father's >> surprise birthday party next Sunday. >> So, the video shows a foldable tablet, a touchscreen interface, [clears throat] a conversational voice assistant with a really strong personality, access to both global and personal information, real-time video generation, real-time computer vision, seamless video call integration, delegation of complex tasks for autonomous execution, and what we might call today continual learning.
And it's really, really clearly influenced by AS we may think, but it's also quite different. It really is updated for 40 years of progress, and it really does sort of presage this AI agent era we're in now in a way that Vannevar Bush's Memex didn't and maybe couldn't. The Knowledge Navigator video divides our timeline, I think, quite neatly in half, and hold that thought cuz we're going to come back to it. The 1990s were about the network, first local area networks and dial-up, and then the internet and the web, and with the benefit of hindsight, I now think that the single most important thing about the web was that it was multimodal from the very beginning.
More even than the GUIs of the 1980s, the web anticipated that text and audio and video and data were not different things to be used in different programs, they belonged together. And in a real sense, the web was an attempt, and a conscious attempt on the part of a lot of people building the web to make that Knowledge Navigator video real. Then in the first decade of the new millennium, the big job was to make all of this computing stuff mobile and continually connected, to put this new multimodal networked computer in your pocket, literally, to give a supercomputer to everybody in the world that they could carry around in their hand.
And as with the 1960s, there was an efflorescence of like futurism on screen in the first few years of the new millennium. And I think it was because computers you could carry around with you and cameras everywhere and a kind of Moore's law for pixels making screens super cheap really gave us a chance to think through what we thought the future would look like in a new way. A lot of stuff we could almost but not quite build was cohering in the minds of people working on these machines.
And the best and most famous Hollywood computers from that era were created by John Underkoffler [clears throat] for the films Minority Report and Iron Man. Here's Minority Report from 2002. >> It's no longer there. >> Time frame? >> 13 minutes. >> Hey Chief, investigator from the Feds here. >> Yeah, I don't need some twink from the Fed poking around right now. >> John, I wrote it down on your calendar. I left you a message at your house. >> Check in with the Favors ahead of Ford and see if the neighbors knew where they went.
Check all relations. >> Check the neighbors and relations. >> But John >> It's actually >> Just get him some coffee. Tell him some stories how I save [music] your ass every day and you can't do without me. >> I got coffee. Thank you. >> Danny Witwer. Twink from the Fed. Oops. Gone. >> So the gestural interface in Minority Report was implemented on screen as special effects, but it was actually based on John's PhD work at the MIT Media Lab.
In a real sense, this was real technology. John had brought the UI out of the small screen and into the world with us in a bunch of really interesting and lovely ways. John also consulted on Iron Man, which is a very different view of the future than you Minority Report, which was Spielberg working in like the American Kubrick dystopian tradition. Iron Man is really squarely in that Star Trek goofy futurist tradition.
But I think you could see the common elements in the UI depicted on screen. It's still from the same era. >> Wake up, Daddy Sean. >> Welcome home, sir. [music] Congratulations on the opening ceremonies. They were such a success. As was your summit hearing. >> [music] >> And may I say how refreshing it is to finally you in a video with your clothing on, sir. You! I swear to god, I'll dismantle you or short your motherboard.
I'll turn you into a wine rack. >> I co-founded a startup with John in 2006 to make the Minority Report interface into a commercial product. This is our demo reel from 2012, 6 years into that work. >> [music] [music] >> This long project to build the multi-modal, multi-device, multi-screen, multi-player, ubiquitously connected computer is still what I'm working on 15 years later. In 2010s, we built out the cloud, which laid the groundwork for the infrastructure and data centers and data capacity we would need to scale up AI training and inference, which brings us to now.
We're building agents. And we're starting to think about agents plus plus. But I think we can also start to think about the next thing, the AI native software that is to agents what today's internet is to the web pages of 1995. And one way to think about the story is this. We went from the calculator to the computer to the personal computer to the global cloud computer. And now we actually have the ability to build the Memex and Jarvis from Iron Man and knowledge navigator from that 1987 video for real, completely working.
And by building those things, we'll figure out what we want to build next. A couple of weeks ago the team at Tavis released a reimagined knowledge navigator video this time entirely built on real and available technology. I'll just play another 20 seconds of this, but like the original knowledge navigator video, this is worth tracking down and watching in full. >> Good evening, Hassaan. I adore that houndstooth jacket you're wearing today.
Anything I can help with or would you like to review tomorrow's schedule? >> Thanks so much, Tom. Yeah, let's review tomorrow's schedule and see how busy it is. >> Opening your calendar now. Here is the quick version since it is late. Tomorrow morning is slammed. Investor meeting at 9:00, then back-to-back one-on-ones and internal meetings until 1:00. >> So, the full 4-minute video is one take, completely real. And when you watch it, it really does feel both like the knowledge navigator video from 1987, familiar but built on real technology, and like something brand new.
And I'll just close with a massively multiplayer game project I've been working on with some friends as a canvas to really think about what AI native software can be. This game is built from the ground up with LLMs as the core of every interaction. At every moment in the game, there are hundreds of inference calls happening and we couldn't have built anything like this even a year ago. Oh, sorry. >> Welcome to Gradient Bang, a multiplayer game [music] that showcases real-time agent orchestration.
Gradient Bang demonstrates several patterns for AI sub agents such as asynchronous non-blocking context compression. >> Okay, make a note for later. We are going to eliminate Heliotrope from existence. >> Noted. >> Long-running sub agents that share context. [music] >> Eagle is on five trade loops. Hawk and Raptor are on five exploration loops [music] each. Your fleet is busy. >> Progressive skills loading. >> How much does your average ship cost? >> Ships range quite [music] a bit, Captain. >> Dynamic user interface generation. >> Show my task history. >> Certainly. >> Uh hide the map. >> Okay. >> And conversational [music] voice. >> No, I don't want to exchange it.
I just want to sell it for cold, hard cash, please. >> I'm afraid [music] the galaxy doesn't allow you to be shipless and hitchhike. >> So, I went long after wrap-up, but I will say that the first version of this new draft talk was an hour. So, I have a lot more things I'm super excited to talk about with all of you. So, if you are interested in this stuff, come find me. We have a booth on the show floor. I'm online everywhere and I'm excited to build agents cuz agents are awesome, but also to build the next next thing, too.
Thank you. >> [applause] [music] >> Mhm.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.