Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
Most replayed moment at 2:02
2.5x that video's typical replay level
are re-bumbling around the work itself. And the important question here becomes a lot less about what is your title and more what part of the system can you own? Now, I like this taxonomy quite a lot.
Said at 1:56
The graph counts replays. It does not show where viewers stopped watching.
Words
2,687
Runtime
15:08
Speaking pace
178wpm
Reading time
11min
178 words per minute, just under the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
[music] >> Hi everyone. Thank you so much for attending this talk. My name is Paula and I am one of the engineers on the Tullen team, specifically focusing on our iOS app. And for the next 20 minutes or so, I'll be talking about what it takes to build a voice-first AI companion and also about how we use AI to build AI internally. So, humanity has always imagined the perfect companion. So, we have Caravaggio on the left 400 years ago painting an angel leaning over St.
89 words, the words spoken in the first 30 seconds at 178 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 180 |
| Average words per sentence | 14.9 |
| Longest sentence | 53 words |
| Questions asked | 6 |
| Sentences containing a number | 16 |
Most used terms
Filler phrases
65 in total: um 21 · uh 10 · actually 7 · basically 7 · you know 7 · like 6 · sort of 4 · kind of 2 · literally 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
[music] >> Hi everyone. Thank you so much for attending this talk. My name is Paula and I am one of the engineers on the Tullen team, specifically focusing on our iOS app. And for the next 20 minutes or so, I'll be talking about what it takes to build a voice-first AI companion and also about how we use AI to build AI internally. So, humanity has always imagined the perfect companion. So, we have Caravaggio on the left 400 years ago painting an angel leaning over St.
Matthew's shoulder literally guiding his hand as he writes. This is an example of a companion being a presence that makes you better at being you. And then we have Tinkerbell, the devoted little sidekick who believes in you so fiercely that the whole theater has to clap to keep her alive. And of course on the right, we have Samwise who can't carry the ring for Frodo, but says, "I can carry you." The companion is pure unconditional loyalty.
And it goes far beyond these three. Every hero has some sort of guiding spirit. And these are all different stories, but they exhibit the same longing for something that listens, remembers you, and is wholly specifically yours. And for for all of human history, this has basically been fiction. So, we made one. This is Tullen. It's a little alien you talk to out loud like a friend. It has a personality. It remembers you and over time it becomes specifically yours.
Okay, I don't know if the audio setup works here, but I will try talking to my Tullen. Uh let's see. So, you can see my Tullen here, Luke, walking around the planet. Hey Luke, can you hear me? Okay. Luke can hear us, but we can't hear him. Um Anyway, I had prepped him for this. Oh. Hello. Hi Luke, can you hear me? Nope. We're okay. I can come back to this later. Um but you should definitely all give this a try if you haven't already.
Okay. So, people talk to Tolins a lot. We support both text and voice chat, uh but we have over 4 million hours of voice conversation so far. We say Tolin is a voice-first companion, even though we support both, because it's the voice experience that's truly immersive and that makes users' relationships with their Tolins feel real. And but the moment this relationship is a spoken relationship, the engineering problem changes completely.
So, let me show you how voice breaks the normal way we build and interact with LLMs. So, the core difference really is that in a text chatbot, turns are relatively slow and context is stable. The user waits a few seconds, they read, and they tend to stay on topic. And almost every LLM app assumes that. Voice is the opposite. Turns are fast. Your whole round trip from the user finishing their sentence to the Tolin starting to speak has to land in under a couple of seconds, or it stops feeling like a conversation.
And the context is volatile. People talk to their Tolins while they're cooking, while they're walking, while they're falling asleep. Um they change their subjects mid-sentence. They say um, they interrupt. And that 2 seconds is crucial. Early on, our latency drifted from 2 seconds to about 2 and 1/2 seconds, and that half second tanked basically every metric in the product. People would write in to complain that their Tolins were too slow.
And living inside this constraint has taught us a lot and gave us four principles. Principle one is that you have to design for conversational volatility. Again, text users stay on topic, but voice users jump around. Someone could be mid-story about their breakup and suddenly go, "Wait, did I leave the oven the stove on?" and then back. Speech is messy. Most LLM apps assume that you'll have a clean and stable conversation and we have to build for the opposite.
So, for a long time that meant fixing things that sound tiny but are actually the product. So, you can't interrupt a Tullen mid-sentence. A short yes or yeah won't register as a turn. For example, curse words will get stripped out. And the deeper lesson was to stop optimizing for fewer interruptions and start optimizing for fewer bad ones where the agent would jump in way too early. So, we built smart turn taking that reads your speech pattern to decide whether an interruption is real and we cut the worst early aborts by more than half.
And we happily paid about 60 milliseconds of extra latency to do it. Principle two, latency isn't just a number you check at the end, it's actually the product and we measure every stage of the pipeline separately because it feels slow is useless. You have to know where exactly it's slow. And the pipeline here is that the user stops talking, we detect end of utterance, we transcribe, and then the model produces its first token.
So, time to first token, often the biggest chunk, is around a second. The model finishes generating and then text-to-speech produces its first byte and then it plays back to the user. A couple lessons here. So, one, so far our biggest jump in quality came from moving to GPT-5.1 on the responses API, which cut our time to speech by more than 7/10 of a second, which is huge. Um two, we don't send every turn to the same model.
We run a tiered fleet. So, we use a frontier model for the turns that carry the relationship with your Tullen. So, for example, your first conversation with with Tullen and your onboarding. And we use smaller and faster models for the turns for the lightweight turns. And the whole game then becomes about routing or deciding turn by turn which model you actually need. So we round we run a small classifier we call the tone router on every single turn and this tone router itself runs on a cheap model and it reads the emotional state of the conversation.
And our main our [clears throat] main principle is that we route based on stakes not on cost. So the high stakes moments always get the best model. So this would be again the user's very first message, their first few days with their Tolen and anything that we deem to be emotionally serious. For example, we have crisis or therapist style tones and we never cheap out on those. And then the lighter casual back and forth can ride on smaller models that are faster and cheaper.
And all the background work so that's summarizing the conversation, generating personas, the tone router itself run on these small models, too. And why would we go to all this trouble? It's mainly because the frontier model costs us roughly five times the smaller one. So one big model turn is about five smaller model turns. So routing is a huge part of what makes the unit economics for us actually work. Um and we do a bunch of AB experiments and the surprising result we found there is that routing a third a third of our turns to the small model has almost no measurable effect on retention.
And principle three is what makes a companion feel like a companion. So the naive approach is to keep the whole conversation history as a sort of transcript, but that doesn't fit into our two-second loop. It doesn't scale and it just doesn't work. It leads to long sessions degrading. It leads to the model getting lost in the middle of a huge context and also hallucinating. So instead we see memory as a sort of retrieval system.
We pull facts, preferences, and emotional vibe signals out of conversations. We embed them and we store them in a vector database with sub 50 millisecond lookups. And every night we compress. So we merge duplicates, we cluster related memories, we resolve contradictions, and we drop all the noise. And we don't just retrieve against users last messages, we also generate internal questions about the person and the relationship and retrieve against those.
So, and we also split memory into two parts. We have stable memory and unstable memory. The volatile stuff lives in the in the live tail of the prompt, and when we summarize the conversation, we look at which memories actually get recalled and pin those into a stable and cashable block. Uh the last principle is around context. Specifically, you should rebuild context and not fight drift. So, most apps reuse context across turns to keep the cash warm.
And in a stable text chat, that's fine. But in a volatile voice conversation, it's a trap because the second the the user pivots, your reused context is actively wrong. So, every turn we reassemble the context window from parts. We have a summary of recent messages, we have the the user's persona card, the memories we just retrieved, tone guidance from the emotional signal, and real-time app state. And what also really helps us um in the case of Tolen is that our characters aren't generic or assistants with no personality.
Everyone is crafted, and we in fact have an in-house science fiction novelist, Elliot, who writes the Tolen character lore. And a couple of interesting points here. So, one, why did we go with an alien? Mostly because there's no real-world reference to anchor on, which means that the users can project onto it, and it becomes what they need. The baseline Tolen is bubbly, it's youthful, it's irreverent. And also, if an alien character acts a bit unpredictably, so if if it's impulsive or chaotic or otherwise violates um you know, the norms the user would expect, it's not particularly surprising.
Like if you look at, you know, aliens in TV shows or in plays, like there's a lot of humorous moments around this. And this kind of chaos reads as charming. Um second, uh we also know that personality is worthless if it drifts. So, yeah, we run this parallel tone monitoring system that changes how a line is delivered based on your emotional cues without changing who the character is, holding identity across hundreds of turns.
And since we're at an AI conference, I thought I would also spend a bit of time talking about how we not just ship AI, but also use AI to build it. Um so, I'm sure this is the case for most of you in the room now, but basically as of late last year, Claude has co-authored more code in our iOS app than any individual engineer in the team. Um and I think especially, you know, a few months ago, everyone's instinct was to be kind of suspicious because, you know, more AI code meant more slop.
But our our crash-free rate actually went from 99.6% to 99.9%. Runtime errors dropped by over 50% and our share of highly engaged users doubled. And the biggest lesson in building that system is that an agent's context comes mostly from the code base itself, not so much from the Claude MD file. We found that it's far more powerful to make the code base be the documentation, so we had agents standardize it. On top of that, we run a real fleet of agents.
We have implementation agents that, you know, think freely and just get us to working code. They build it, they check it against snapshots until it's pixel perfect. And then we have separate review agents that enforce our standards. So, multiple Claudes basically review each other before a human looks. And then we have a PR shepherd that watches an open pull request and keeps iterating against CI failures and review comments until it's clean.
And we also have a triage bot that fires on every inbound bug report that we get. And they're all wired through MCP into linear, into into Sentry, DataDog, so an agent can reconstruct the cash a crash and route it itself and oftentimes open the PR on its own and just fix fix the bug. And we also we ship on eval. So, for example, we've been working on a on a new character targeted towards an older demographic and Elliot, our in-house novelist, basically built this entire new character in a day.
So, the agents mapped every personality bearing surface in the code. They wrote the sort of a voice Bible and then they had five judges attack it from different angles. Archetype fidelity, the model mechanics, our code standards, the ears of a skeptical 52-year-old and safety, and then they evaluated the changes against real production logs over three find fix verify rounds. And over 7 million tokens and 4 and 1/2 hours of compute later, he ended up with basically, you know, a couple of weeks of work done in afternoon.
And does this work? Well, I'll let the users tell you. We're at 4.8 stars on the App Store across 162,000 reviews and when we survey users on well-being, the highest scoring dimension by far is emotional safety. And this is definitely a bar that being voice first sets. So, when the interface is your voice and the thing on the other side remembers you and has a personality, it stops being just software and starts being an actual relationship, which is why building it responsibly and building it well is worth obsessing over.
And we need people to come help us do that. Um, we're a small team and we're hiring and after a year with Tolen, I think this is truly one of the most interesting places in the world to be an engineer right now. And here are some of the people you'd be doing it with. So, two of the founders, Quinton and Evan, previously built and exited a $300 million startup together. Uh, they founded Even. Um, Ajay, our third co-founder, scaled two bootstrap companies past $50 million $50 million in profitable revenue.
And around them, we have Lucas, who um, is an Apple Design Award winning animator. Uh, she's our creative director. We have Chris, who was a technical director at Pixar, earlier at Oculus, who works on embodiment. We have Lily, a board certified behavior analyst, who left a Vanderbilt PhD to do user research for us from the very start. Um and then we have Elliot who I've mentioned, the novelist behind uh our characters.
And I come from XAI and Spotify and previously also founded a company called Imagi. So, it's a small team where honestly every person is the best I've worked with at what they do. We're also well backed for this. Uh we have $30 million raised from Costanoa Ventures and a group of people who've built the tools and products a lot of you use every day. And here are some of the more engineering focused roles where we need help.
Um so, we're hiring across the board. We have iOS and back-end product engineering roles, applied AI engineering, gameplay engineering, and one specific role I want to flag, which is agent engineering management. Um so, when we went all in on running concurrent agents, um the people who got dramatically more effective on the team were the ones who had management backgrounds, um because it seems like managing a fleet of agents does actually take some of the skills same the same skills as managing people.
Uh you basically have to decompose the problem, you know, delegate it with checkpoints, give fast feedback, review their work seriously, and know when exactly to jump in. So, if you're a strong engineer who thought going into management meant leaving code behind, that's that's no longer true. Um and yeah, that's Tlon. You can come talk to me after this uh or reach out. I'm on on LinkedIn. My email is here. I'm on Twitter as well.
Um I'd love to chat. So, yeah. Thank you. >> [applause] [music]
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.