Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
Most replayed moment at 2:02
2.5x that video's typical replay level
are re-bumbling around the work itself. And the important question here becomes a lot less about what is your title and more what part of the system can you own? Now, I like this taxonomy quite a lot.
Said at 1:56
The graph counts replays. It does not show where viewers stopped watching.
Words
3,504
Runtime
20:27
Speaking pace
171wpm
Reading time
15min
171 words per minute, between the 160 25th percentile and the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
[music] >> Thanks so much for having me and inviting me to to be a speaker at the the welfare. You know, I I attended last year and was so impressed by the quality of presenters. So so glad to to have a chance to be here and to present. So the title of my talk is, you know, video has no memory, right? And you know, this might sound strange because video is already in like a preservation of the past, right? If you think
86 words, the words spoken in the first 30 seconds at 171 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 228 |
| Average words per sentence | 15.4 |
| Longest sentence | 51 words |
| Questions asked | 62 |
| Sentences containing a number | 4 |
Most used terms
Filler phrases
319 in total: uh 76 · like 64 · you know 55 · right? 53 · um 50 · actually 10 · basically 4 · kind of 3 · sort of 3 · I mean 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
[music] >> Thanks so much for having me and inviting me to to be a speaker at the the welfare. You know, I I attended last year and was so impressed by the quality of presenters. So so glad to to have a chance to be here and to present. So the title of my talk is, you know, video has no memory, right? And you know, this might sound strange because video is already in like a preservation of the past, right? If you think about like you have footage, you preserve recording training data, incident, creative work, history, etc.
But actually most of the video AI systems these days do not have memory in the system sense. So actually for this talk, I will try to answer the question like what could it take to build a memory layer for video intelligence? To start, I want to be clear about what makes video different from other data type, right? So this is the first mental model that I want to highlight, which is that video is not a stack of frames.
Um so in in a many of my conversations with like developers, you know, who are using our product, a lot of them still treat video as like a stack of images, maybe a transcript being attached or you know, like you know, but essentially like like a frame level, right? And that is a useful approximation for some tasks, but it throw away the thing that makes video very unique, which is continuity, right? So meaning in video derives from space, time, modalities, and sequence.
So a better mental model for video is a special temporal volume. So what I mean that inside that volume, you have visual information, speech, sound, motion, OCR, camera changes, scene transition, metadata, and time, right? So the hard part here is really well, how can you preserve any relationship across this volume so that later an application can traverse it? And then you know, especially at the enterprise scale, you know, across like industry like entertainment, sport, you know, um, short-form content, then you're sitting on petabytes of of footage, right?
So, finding moment is already hard. So, how can you preset meaning across millions of moments in the deeper platform? So, I work at Twelve Labs, which is a uh series B uh startup. Um, we build foundation models that understand, you know, video the way that human do. Um, and the way we talk about our positioning is like the existing uh stack of of dealing with video is not equipped to do that, right? Obviously, language model are very powerful.
They are good reasoning interfaces. They are increasingly multimodal, as well. But, the supporting stack, right, around that is very um I'd say limited, and that create three problem. Number one, is wrong context, right? So, video is not naturally a sequence of text token. If we force it into that sequence by sampling frames, by extracting a transcript, uh by dumping everything into a prompt, you lose the spatiotemporal relationships, right, that actually define the event.
Second, uh wrong memory. So, if you think about text system memory here is often mean retrieval and generation, vector search, or probably like larger context window. Uh those are very useful, but video memory has a different requirement. It needs to link today's scene for something that happened in another file, another episode, another camera angle, another season, another year. So, it actually need durable continuity, And the last part here is wrong reasoning.
As I said, you know, text-first system cannot reason over, you know, um, natively over motion, causality, all that. So, uh you know, they do not automatically build like a persistent structure uh on, you know, who appear, what happened, what changes, etc. And so, my argument is that video intelligence need a memory layer that decide what to preserve, how to connect it, and how to retrieve later. So, um, I want to kind of ground it into the properties of video, right, to to make it even clearer.
There's five challenges in dealing with video. Number one is temporal, right? So, depends on before and after. So, frame by itself can can be misleading, right? The same expression, product shot, physical action can mean different things depending on the sequence around it, right? Second is that video is obviously multimodal as I I explained already. Um you know, a transcript alone may miss, you know, the logo, a frame alone may miss the spoken claim.
Video Video is also very dense, right? So, a few minutes can contain dozens of shots, people, objects, action, location claims. The useful signal is uneven across the distribution on on the frame. Some seconds are decisive, others are noisy. Fourth is that uh video is also ambiguous, right? Um people reappear under different lighting and angles, brands are partially visible, location are implied, concepts emerge over time rather than being named in a single moment.
And lastly, it is uh expensive uh because in a lot of uh big enterprise and in complex workflow, you need to, you know, point back to the source moment, like where it come from, right? So, these are the five properties explaining why video memory um is is very complex. You need to preserve temporal span, multimodal evidence, continuity, all of that. Uh this is a very simple uh stack of how we build things at Twelve Labs.
Um at the bottom, we have these semantic chunks that capture, you know, meaningful temporal units. Above that is our uh multimodal embedding coder called Marengo, which essentially turn those spans into spatial-temporal relations. Um basically, vector embeddings that represent video content. And then we have a special special-temporal context store, which is where it preserve pre-usable structure, like moment, entities, metadata, all of that.
Uh we also build our own uh VLM, video context-aware language model called Pegasus, that essentially serve as the the reasoning layer, right? That can preserve uh prepare over video content, so think about summaries, like um metadata, synthesis, comparison. And we expose our our models as API because, you know, we want to get developers to use them as infrastructure. Now, moving beyond like kind of the stack right here, I want to talk about the difference between search and memory, right?
Very quickly speaking, uh search is obviously super important. It's how you recover relevant moments from large video library. But then it gives candidate. It actually not give you like any continuity. So, memory on the other hand is is all the um you know, uh the things that enable the system to answer different class of question as you see here on the right side of of my screen. So, these are not the single retrieval call, right?
They require the system to reason entities, timeline, evidence across an entire corpus. Um and so like you can actually build product moving beyond from like show me something like this to, you know, tell me what this collection knows, right? And so that that might sound, you know, simple and subtle, but uh the the the the output is completely different. Like with search, you you get like an output like a time-bounded moment, but with memory, you actually return like structured knowledge, timeline, uh explanation, composable output.
And that like, you know, is very important because we can now move um the the unit output from clip retrieval to corpus memory, right? Um there are two scaling dimensions shown here in the slide. The first is time scaling. So, a real video system should be able to reason over years of footage without reprocessing the whole archive every time, right? That means memory first retrieval, uh be reusable representation once and then support multi-hop timeline, episodic recall, follow-up question at lower latency and cost.
And then the second uh dimension is in space, right? So, many real workflow actually um involve multiple perspective like different camera angles, uh you know, live stream, creator broadcasting content, body cams, stock cameras, uh event feed, right? So, how can you build a system that can fuse evidence across all the sources and then maintain current understanding, right? And so, that is the challenges here. How can you build a representation that let application traverse video across time and across sources?
Um since this is um you know a track on on graph, right? So the the best mental model that I can come up with is to represent you know, video collection as a context graph. So a context graph is a durable, queryable representation that connects video moment, entities, appearances, relationship, time span, metadata, and corpus level context, right? So if you take a look here on on the screen all the way in the bottom, you got time bounded moment.
These are like the the scene, the shot, right? Uh these are evidence unit. One one level up are the appearances, where and when each entity show up. And then you got the actual entity itself. So think about the people on the video, the brand, the places, the concept. Next you have relationship, uh co-occurrences to the same brand, sequences between different places, the causality, and timeline. And finally at the top you have corpus level context.
What are the main themes, the patterns, the gap, the coverage that this video collection cover, right? Uh this matter because different question travels different part of the graph. If you ask a simple search question, then that might go directly into the moment. But like an entity workflow might start with a person and then it expand into appearances, right? And if you ask question like a storyline, like narrative storytelling of certain you know you know person, then it may follow relationship across time, right?
Uh so the key idea here is that memory in the context of video understanding is a navigable structure over the entire video volume. From that concept, I come up with these five principles as a building you know a memory layer for video intelligence. Number one is to ingest once and reason many times. So um you don't want to like do every single query from scratch, right? You want to pay the cost up front, do one interpretation from the video content up up front, pay the cost, and then you move expensive understanding into injection.
So this is same mental model uh database, right? You you're not repeatedly have to parse your entire source of data um for you know every application request. Um second principle is to store primitive, not just answer. So, uh, you know, moments, entities, appearances, as we talked about that, those are the the primitives, right? That allows you to, uh, do downstream workflow, like search, editing, um, you know, analytics, all of that.
Third is to ground every claim. They like basically, if you ask a question, you need to cite back into where that scene happening in the video. So, uh, evidence, like, you know, should be grounded to a specific timestamp within the video, right? Uh, fourth is to let intent shape memory. Um, this is important because the same footage mean different thing in different workflow. We work across sports, uh, application, brand safety, compliance review, clear analytics.
All of them require different primitives from the same video. So, the memory layer should be configurable, right? Developers should, uh, should be able to tell the system what matters. And lastly, uh, keep the layer composable. Um, so, basically, being API first, you know, um, it should provide, uh, the layers that allows those application, uh, on top of that to to serve it, uh, structured grounded metadata that can be plugged into any sort of application.
Um, so, moving beyond these five principles, I want to talk about like, kind of the the harnesses around building a memory layer, right? There's a lot of talk these days about, um, you know, building the right harnesses for the context of language model. So, what does it look like for for video understanding model, right? Um, a model could produce a single answer. It is stateless. It start fresh each time, start fresh each time, and doesn't have any constraint.
So, the output is largely based on what the model is set to produce. A video worker, on the other hand, operate inside a a very deterministic system. Understand what is available, uh, it can plan the task, could ship evidence, inspect, uh, the relevant moments, synthesize, validate, return output, and then the entire workflow can be evaluated, right? Um, so, so for video understanding, this is very important because the worker need to know what memory is is available, what evidence matters, and also like how deep to inspect, because that will depend, uh, determine how much cost to spend, what output constraints to satisfy, right?
Um talking about harness engineering for for video understanding, um I come up with these like different capabilities for for like a video worker, right? Um number one is memory, I talked about that already. Number two is task planning. So, given given a query from from an user, uh you have to decide like what task to execute. Is it like search or is it like summarization or like, you know, multi multi-step reasoning?
Uh third is retrieval, like every single system should be able to like select the right evidence from from your video corpus to read for a specific task. Uh expert tools, right? So, we work with customer where they require like, you know, zoom in, zoom out, uh comparing different uh uh you know, frames, enriching uh content with like additional metadata. So, building expert tools inside uh like a like a video worker uh is very important.
Operating envelope, so these are like explicit limit on time, cost, depth, scope, autonomy. Uh and output contract. So, sometimes natural language is not, sometimes the patient needs structured data with references and time stamp. And of course, finally you have evaluation, right? Uh like, you know, did the retrieval find the right evidence? Did the synthesis preserve it, right? Did the worker stay within the budget?
All right, so so that's a lot of like, you know, uh slide and and and talk. I want to quickly jump into some demos uh that I actually built using Two Labs uh you know, video agent product. So, the uh there'll be three demos. Um the the video agent product that we've been building is called Jockey. So, this first example here is for sport understanding. Uh you know, obviously everyone is super excited about the World Cup that happening right now.
So, what I did is I ingest um 67 videos from the 2022 uh World Cup in Qatar. And you can see here I asked it to find the near misses, uh the shot that almost become goal but did not. For each, explain why it was not a goal, uh but do not include actual goals, right? So, these are the the the top output that it returned. So, that that's hitting the woodwork. This is saved from the goalkeeper. I don't know if the sound is up, but like I'm playing the the video, by the way.
Um Right, this is another save from the goalkeeper. It even catch like, you know, the upside uh from one of the goals. And then I ask question, okay, when is a goal? Uh Find the most dramatic actual goals. Show the build-up play and the finish. For each goal, describe the sequence, right? So, if you know this one, this is the um the first goal of the World Cup final like 4 years ago. And it it actually like returned like, you know, you you don't understand who who are the passer like it was Alvarez passing to Mac Allister and passing to Di Maria to score the goal.
Um Take a look at this one from Richarlison. This is goal of the tournament uh from Brazil again uh South Korea, I believe, right? And it returned like, you know, um an outrageous skill in a build-up, right? The name the player who did the return pass. You can even do player tracking. So, I asked it to track Lionel Messi across this entire corpus including the shot where he's won the man of the match figure on the screen.
Describe the camera framing, right? Uh so, this is uh hilarious of all the important moment in the game and this is a scene where Messi dribble past a sliding defender. You can see here. It pick up the scene where he scored the first goal against Australia in the round of 16, I believe. Right? This is another scene where he scored the third goal in the final. >> Yeah, so that one example on spot spot understanding. But then you can obviously build more interesting and more like real practical application of which security is one that we encounter a lot.
So on the on this example I ingest it you know publicly available camera footage. For context this are the clip you have traffic jams suburban you know urban area and given this footage where I asked Jockey to count and classify every vehicle in the intersection break it down by type plus pedestrian. And it returned the number of vehicles and the big foot traffic as well. You can detect safety events right? So you see that a red SUV turn and almost gets struck.
Another scene here. Turn left into an upcoming car. Yeah, so that a clear straight line entry. It works well in you know different scenario. This scene is very crowded area in Bangkok. It asked it I also asked it to work on the you know the rain right? So this is another scene where it understanding the um rainy condition. It identify the basis intersection vehicle window. So yeah, those are some example for for camera security surveillance footage.
Finally advertising so um you probably seen this Adidas clip in all commercial leading up to the World Cup. Recently is a 5 minutes Adidas footage and I asked it to classify all the point where you can put an ad on. So, it find the review, the hard cut, the impact, energy big rate. It find the scene where a certain player appear on the screen. It identify like, you know, high impact action like this. Condition, um, the hard cut to Nike football underlies.
And of course, it it point into the logo. Um, of Adidas. So, you know, uh, from from perspective of an advertiser, these are very important moments because they can, you know, find the scene with the slow motion hero. Hard cut on a beat. Or pick action. In which they can advertise their brand content against this footage, right? Um, yeah, so those are three sample demo application, um, that I want to highlight, uh, I've using two labs.
Um, and again, um, you know, what can you build with with this sort of video memory layer based on the case example? These are the categories of of application that I believe developers can build. You can discover things. You can be reasoning experience. You can organize your content across different video library and it can be action workflow, assemble uh, different scene together, do compliance review, data operation, etc.
The same framework apply for different verticals in media and entertainment spot, segmentation, highlight generation, in commercial security, evidence review, contextual analysis. In advertising, uh, brand safety, uh, creative intelligence, right? And, uh, yeah, so this is our product, uh, that I coming up. Um, one quick highlight is that we we try to as a video cognition infrastructure. So we have a knowledge style that basically become the video memory layer.
Web configurable injection that let video shape what the system can can extract. Corpus digest so that you can understand in what is in the library and the resolution as you search responses API. So the thing we want to highlight here is it's not an application layer. It's not an editing platform, not a compliance product. It's the cognition infrastructure with the layer and the harnesses that enable like those product being to become available.
Um and if you found the content of this talk interesting, definitely recommend you to scan this QR code. The product is currently in private beta right now. If you bring any sort of workflow that touch video content especially around content assembly, content organization, you know, think about media archive, content creator, YouTube, TikTok, sport analysis, media workflow, definitely add the scan this QR code and register for the interest or come talk to me after the talk.
So that should be my time. Thanks a lot.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.