Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
Most replayed moment at 2:02
2.5x that video's typical replay level
are re-bumbling around the work itself. And the important question here becomes a lot less about what is your title and more what part of the system can you own? Now, I like this taxonomy quite a lot.
Said at 1:56
The graph counts replays. It does not show where viewers stopped watching.
Words
3,221
Runtime
15:50
Speaking pace
203wpm
Reading time
13min
203 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
[music] So, my name is Charlie uh and I work on the developer experience team at OpenAI. And part of my job is talking to developers to understand and see, you know, what and how they're building with our models. um whether that's text, image or audio. And lately I've been thinking about a misconception that I have seen or maybe it's just a misunderstanding and it's the idea that voice agents have to talk back. >> And to some of you that might sound, you know, absurd. It's a voice agent. What do you mean it's not supposed to talk? Uh,
102 words, the words spoken in the first 30 seconds at 203 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 147 |
| Average words per sentence | 21.9 |
| Longest sentence | 126 words |
| Questions asked | 42 |
| Sentences containing a number | 5 |
Most used terms
Filler phrases
288 in total: um 91 · you know 56 · like 50 · uh 45 · right? 29 · sort of 7 · actually 5 · kind of 4 · I mean 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
[music] So, my name is Charlie uh and I work on the developer experience team at OpenAI. And part of my job is talking to developers to understand and see, you know, what and how they're building with our models. um whether that's text, image or audio. And lately I've been thinking about a misconception that I have seen or maybe it's just a misunderstanding and it's the idea that voice agents have to talk back. >> And to some of you that might sound, you know, absurd.
It's a voice agent. What do you mean it's not supposed to talk? Uh, but I think if there's one thing that you take away from this presentation, I would like it to be the idea that speech is not the only way that a voice model has to respond. And I think models are getting intelligent enough and capable enough that they're starting to open up some uh new modes of design. I mean there's actually three kind of modes that that you know I kind of see emerging these days right uh speech to speech speech to action and event to speech and there's a couple of things I think worth pointing out about um these three categories.
The first is that they're not new, right? I think as we've seen from from previous talks, even just today, um there's a long history of building these types of systems in and around voice. Um if you squint, you could make the argument that the movie phone hotline where you called in to get showtimes was an example of a speech-to-pech system. Uh and I think you could pretty reasonably make the argument that you know GPS navigation in your car which has existed since I was a kid is an example of an event to speech system.
So, it's not that they are brand new, but I think it is that we are able to do some some much more interesting things with them uh now that uh we're in this era, right? And I think, you know, the other thing I would mention here is that um they're they're remixable, right? They're not meant to be mutually exclusive. Um and I think as we'll see in a little bit, the best products uh exist in a way that combines all of these modes.
So, speechto um everybody knows it. Hopefully, everybody loves it. the user talks uh and then the model talks back and I think there are a few examples that that I can give um for this type of use case right you've got things like live practice and coaching especially around language learning right I think the ability to um hear a lot of emphasis or emotion and give people that feedback is really powerful um I think you can have you know what I am sort of cheekily calling concierge experiences which I think is just another way to say customer support plus+ Um the first voice tutorial that most people you know try to build when they have access to this technology is some sort of customer support chatbot for very good reasons.
But I think with you know when we add a richness and a depth to the voice models as has been happening in recent months in recent years um we can build something that is like much more enjoyable to use than like talking your way through a phone tree. Um, and so, you know, I have the hope that like soon, if not like, you know, now, we're capable of building support experiences with agents that actually feel much more enjoyable to talk to than like arguably like the median human support agent.
Um, and as we just saw if you hear the last talk, uh, live translation, right? The models have gotten good enough and fast enough that we can just dynamically translate content on the fly, um, with like little to no latency. It wouldn't shock me if at next year's keynote um you know they live streamed it from the main stage but also uh dubbed it in real time across multiple languages. The second category is speech to action.
Uh these are talks and the model uses tools and I think this is one of the most underexplored areas that we have. Um I actually almost titled this talk uh voice is the next capability overhang because I think there is just a vast vast amount of stuff um that we could be doing in this category that we are not currently doing. Uh for example um there's a broad spectrum I don't have you know there's way too many examples even fit on this slide but three categories that that I find particularly interesting.
Uh first is form filling right um so much of the internet is just filling out forms. Um, and there is, you know, today no reason why you shouldn't be able to just talk. You know, I would love it if instead of spending an hour filling out a government document, I could just talk for five minutes and it would get 90% of it for me and I would do a quick check, you know, just to make sure that everything looked good, right?
That is a vastly superior experience than like having to type in every single name and address that I've lived in the last 5 years and, you know, all of my previous identities. Um and so I think I think that one is though it may seem boring you know affects a a significant GDP of the internet right the next category is creative tools uh where I am privileged enough that I can speak the language of software and so I can tell codeex you know here's exactly what I want you to build and I can articulate it in a way that um I get much more leverage than sort of just like cludily trying to iterate one thing at a time but I can't do that when it comes to you know using making music or painting um and so if I don't have the ability to articulate um the exact aesthetic that I'm looking for.
Um and if I don't know how to use Photoshop or Ableton, I'm left in this state where, you know, my my taste exceeds my capability. Um and so I'm really looking forward to integrating voice into creative tools so that I can just sort of cludgy go along and, you know, vibe create, vibe compose, vibe paint, um and make something that that's really beautiful to me. And I think the generalizable um category here, right, then just starts to become computer use.
And we've already seen some companies start to do this. Um, you know, it raises the question of like, look, if the models are just getting good enough to do everything on a computer that a human can do, like why am I talking to an app? Why am I talking to a terminal? Why am I not just talking to the entire computer? Uh, and so I think that's sort of a really interesting uh way to start exploring. But if you're a developer today, right?
Whoops. If you're a developer today, um, what does that mean for building your own software, right? And I think it is like much easier than you think to start adding audio as an intelligence layer to the intelligence layer to the apps that you already have. Um, if you're building a modern web application, you already expose so much of it as like action as nouns and verbs, right? And if you think about all the verbs that you have, you have uh API endpoints, you have, you know, React hooks.
Each of those things can like pretty relatively easily be converted into a tool that you expose to a model and then you can give the user the ability to just drive your existing software um with their voice, right? And yes, you still need guardrails, you still need safety checks. Like many of the talks today are going to talk about securing and you know productizing this, but um for this I just want you to think about you know what would it mean to take your existing software and just talk to it.
Um and to go back to that misconception, right? I think there are a lot of um you know like if you're talking to the software maybe it can talk back but we've been developing other ways of communicating with the user for decades right we know these things we know we can show notifications and popups we can change state like the color of a button or a drop shadow we can highlight text um if you've used computer use in the codeex app you know there's this amazing like little ghost cursor animation that goes around and clicks things for you so we don't have to use words to actually tell the user what is happening on screen with their software Uh and then the last bucket here is event to speech, right?
Um the model receives an event and talks to the user. Um and sort of the counterpoint from speech to action. I think this one is still very very exploratory, right? Um you know, if you saw Quinn's talk, I think there's a lot of uh space here of like things we can do. Um and to me, we haven't quite seen what AI native really looks like in this vein yet. But um of the things that I've seen, I think there's a couple of through lines that I tend to notice, right?
The first is hands-free or screen-free experiences. There might be times where uh I need to interact with software, interact with objects and I can't use my hands or more importantly my attention is diverted elsewhere. Um that might be something you know like uh a recipe app. Maybe I'm cooking and I need to just say like what's going on and and have something else have something happen. Um the other category is proactive outreach, right?
Where you the model needs to be able to tell you something or get your attention in a way um that you might not be looking at, right? I think every developer um has uh an endless amount of notifications and events happening in their software. But um no developer in their right mind would sort of say I should show all of these logs. Nor you know would they say I should speak all of these logs. But we can start to conceive of voice as this like upper level in this escalatory path of like okay maybe you animate something and then maybe you pop something up and then if that doesn't work maybe you talk to the user to get their attention.
And underlying both of these categories and I think this this you know this whole presentation is this broader theme of accessibility. Um on a personal personal note, I know like multiple developers who um over the course of their careers lost mobility in their hands, lost dexterity in their fingers and for many of them, they thought their career as a programmer was more or less over. Um and then came large language models, right?
Then came coding agents and voice agents and now they generate orders of magnitude more code than they like previously did um you know on a given given day or month. Um, and so I think there's there's a lot that we can unlock here uh for the broader world as well. Um, to go back to like I said, you know, I think like when it comes to these three modalities, you can mix and match them and we already have some, you know, rudimentary ways that we're seeing this.
I think there's things like, you know, all of these pieces for in-car assistants exist, though nothing has quite like combined them into this seamless way. You can talk to like the CarPlay dashboard. Um, you can tell it, "Hey, go play some Spotify music for me." Um, and then it can come back and tell you, Google Maps can come back and tell you, hey, like, you know, there's traffic on this route. We're going to reroute you.
But like we can now start to think about what does it mean to combine that into like a single voice agent across multiple modes. [snorts] Um, similarly, you know, there's a lot of experimentation in the game space with multimodality. Um, you can think about a real life character where you're talking to it to, you know, mine information about the game, about the world. um you can talk to it to execute actions on your behalf and then it can react to like world events right that are happening and then give that information to you rather than just like a simple notification um and I think the question you know behind the question here right is like I mentioned we've had all these things for a while people have been prototyping them for a while why focus on them now why think about building with them now um and I think that brings me to uh a little bit of context here right as as I'm hopefully most of you know traditionally voice agents are built in this chain Ed model, right?
You uh talk, you transcribe, you send that to a language model. It calls tools. Hopefully, it doesn't take too long to respond. Um it then generates text output. You make that into audio and then you play that back to the user. And some time ago, um OpenAI, you know, decided on a different approach, right? The real-time model family does not do any transcription behind the scenes. It is trained on native audio as tokens.
So, you send audio in and you get audio back out. And the industry I think in general has been, you know, trending more in this direction. Um, and not just making it native audio, but even just letting go of the turnbased abstraction that we've had, right? Um, and so making it that it's just continuous streaming audio in and out. And the reason that opening I did this was, you know, turns out there's a lot of stuff that you lose when you transcribe speech and when you transcribe audio, right?
Um there's the old saying that when humans communicate face to face, 55% of the information is in body language, another 38% is in your tone of voice and the last like 7% is the actual words you are saying. Um and so when you transcribe, you lose tone and cadence and emotional uh you know impact, you lose like whether they're trying to interrupt you, you lose background noise, all of this stuff which is really important context for the model to understand.
Um a much more quantitative reason to do it is that you know the first two uh voice modes in chat GBT were built with this chained approach. Um and as you can see had you know significantly higher latency than using uh the native approach with advanced voice mode. And that brings me to GPT realtime 2. Um and this is going to be the one part of the talk where you know I make my shameless plug. Um real time 2 is the the latest model in the real time family.
We released it a couple of months ago. Um, and the the really cool thing about this model is that it brings reasoning to the audio medium. Um, and so much like our text models, it can now think before it speaks. Um, I'm sure many of us have seen some demos of voice models saying things that are a little bit less than intelligent. Um, and so you can now, you know, try to ensure that you give it more reasoning budget uh to come up with a good answer.
Part of why that's also useful is that we introduced tool calling a little while ago. And so the model in addition to thinking it can also delegate parallel tool calls. Um you can start to bring these together. Um though of course that adds latency, right? And you know that's why we also added preamles. Um preamles are a way that you can prompt the model to uh give the user a heads up if it's going to be thinking or if it's going to be calling tools.
Um you know if you think about the scenario of a travel agent, right? If I called the travel agent on the phone, uh you would want the travel agent to say to say, "Hey, like I'm going to go check flight prices, right? give me a couple seconds to do that. Um, and now with an AI travel agent, um, and preamles, you can actually have it communicate that to the user while it's performing actions in the background. Uh, there's a few other things here, right?
It's got longer context, better domain understanding, um, more natural voices, and it's much more steerable. Um, there's some really cool features, uh, that it can do when it comes to like wake words and just waiting for you to to tell it. You you can give it a name. Uh, you can say like, you know, hey, Marin, do you want to say hi to the room? Um, and if you've like prompted that into the model, then it'll, you know, go ahead and and respond to you, right?
Uh, and of course, you know, uh, obligatory benchmark slide, uh, it does pretty well on the the latest audio benchmarks, too. So, TLDDR, uh, is a pretty good model. Um, but I think the the kind of final thing that you know I want to leave you with here is um when building voice agents uh not to start with the question of like what kind of voice agent am I trying to build, right? I think the thing I want to leave you with is start with the question of like what is the role of voice and audio in this interaction?
Um and then how do I move forward from there, right? Right? And often when I ask that question, it leads to a bunch more questions after that. Things like what can the model perceive? What context does it have? Right? Um what tools are available to it and which of those tools should it be, you know, uh executing safely and correctly? Um should it communicate now? Should it wait? Uh you know, how should it communicate?
Should it be sending visual notifications or using audio? Um, and so taken together, yeah, I hope everybody in here can can start to build some much richer experiences with voice. Um, because like others have said, uh, I do believe that AGI will be spoken, not typed. Thank you very much. I'll be at the OpenAI booth, uh, for any Q&A after. Um, yeah, have a good event. >> [music]
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.