Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
Most replayed moment at 13:49
2.3x that video's typical replay level
Fable uh and it runs into an unknown, ask it to log it, right? So that um you uh you can see where the deviations happened and then you can sort of figure out why as well, you know? It will usually give you some context about what happened.
Said at 13:43
The graph counts replays. It does not show where viewers stopped watching.
Words
4,006
Runtime
21:03
Speaking pace
190wpm
Reading time
17min
190 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
[music] >> I think we can get started. Hey everyone, I'm Jerry, co-founder and CEO of LlamaIndex, and today I'm excited to uh give a talk called building the document context layer for AI agents. Um really uh big shoutout to the AI Engineer World Fair for hosting. Um and if you've seen some of my earlier talks from the previous uh AI Engineer conferences, uh we've kind of traced through a lot of the evolution of how, you know, um agent advances have correlated with uh you know, how you inject context into evolving
95 words, the words spoken in the first 30 seconds at 190 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 161 |
| Average words per sentence | 24.9 |
| Longest sentence | 75 words |
| Questions asked | 4 |
| Sentences containing a number | 17 |
Most used terms
Filler phrases
344 in total: like 78 · um 60 · actually 45 · kind of 45 · you know 45 · uh 38 · basically 23 · sort of 9 · right? 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
[music] >> I think we can get started. Hey everyone, I'm Jerry, co-founder and CEO of LlamaIndex, and today I'm excited to uh give a talk called building the document context layer for AI agents. Um really uh big shoutout to the AI Engineer World Fair for hosting. Um and if you've seen some of my earlier talks from the previous uh AI Engineer conferences, uh we've kind of traced through a lot of the evolution of how, you know, um agent advances have correlated with uh you know, how you inject context into evolving applications.
Um and so today, in 2026, we'll kind of talk about three main topics. One is, you know, what is RAG in 2026? Um two, uh basically that basically decomposes into an agent harness plus a context layer. So kind of going into the modern document context layer for agents, um and giving you a little bit of sense of kind of some of the core capabilities today, um especially uh unlocking, you know, the vast trove of document-based data today, and giving you a glimpse of what's next.
All right. RAG in 2026. But before that, just a little bit about the company, and I'll kind of get on with the main talk. Um you might have seen us as a RAG framework. We started in 2023, got pretty popular, um kind of created a lot of techniques around like advanced RAG, that type of stuff. Um today, we're basically the main document infrastructure for AI agents. Um we deliver the best platform for agents to actually read and operate over documents, and we see ourselves as unlocking basically the vast troves of document context even as agents and models themselves get better.
All right. So, what does RAG actually mean in 2026, and how are agents actually inhaling and addressing context today? This is a snapshot from um actually 3 years ago, um when I first gave a talk on this, and naive RAG was basically just building a simple chatbot over your private corpus of data. If you flash back all the way to January of like 2023, um you know, this technique basically consisted of you have some corpus of documents, um you chunk it up, embed it, and put it into a vector database.
You do some naive top K retrieval, and then you generate some stuff with an LLM. All the steps are fixed. You use kind of like a fixed set of techniques, and then with that you actually get some basic application results by just being able to chat over a corpus of documents. Of course, the agent landscape has changed quite a bit and quite dramatically since then. One is agent loops and tool use have gotten a lot a lot better, especially as both the models and agent harnesses have increased.
And there's basically a cleaner separation now between agent reasoning and how they actually interact with context. If you look at what I call kind of the modern generalized agents, which includes, you know, all your favorite applications and tools out there from Claude Code, Claude Code Work, Open Claw, Codex, and a few others, you know, the retrieval complexity has started to get baked into the agent layer. So, instead of coming up with a variety of hacks to really work around the limitations of naive top K retrieval, you can start getting the agent to really reason about, for instance, the best key best keyword to search for, best like, you know, search term to actually get back good results.
Even if the retrieval tools are basic, the agent can input the right queries to basically loop upon itself and help solve the task at hand. Number two, context is moving up the stack. There's been a lot of conversations, I think, for the first two and a half years of Gen AI of how do you actually manage the agent context window, make sure it doesn't overflow, that type of thing. But, I think as, you know, a lot of the evolving techniques, compaction, long contexts have evolved, I think more and more of the conversation is actually how do you just hook up the right MCP servers and skills and tasks to the agent to enable it to do various types of tasks.
And so, that also applies to agent orchestration, right? Coding agents have gotten better. Abstractions of actually defining what types of tasks and programs you want to build have moved a little bit upwards towards English as opposed to through code. And so, maybe in 2023 to 2025, you define programs and you still build stuff via like importing Python, using TypeScript through code. It's pretty clear these days more and more people are just building stuff using English, whether you're a software engineer or you're a non-technical function like go-to-market marketing.
And you're defining runbooks through English and kind of like defining the right goals and making sure the AI is aligned on the right task. And we see that carrying over going forward as well. So, in terms of just like general evolving knowledge work patterns, and this kind of forms a foundation for what we how we think about the context layer. You know, in the past maybe you could use AI to just like ask simple questions, get it resolved.
Today, you're starting to be able to define and solve more end-to-end like tasks just through English and also use English to start compiling repeatable programs to actually execute something at scale. Obviously, there's been a lot of discussion on how this will evolve even more, and we do see AI agents as behaving even more autonomously through solving long horizon tasks, looping upon themselves to be able to achieve a goal.
So, the future will move towards a state where you actually might not even have to define the task in English, but actually more of the goal and a scoring rubric, and then the agent will use all available context available to it to actually solve the task. So, with that overall framing, let's talk about what kind of what is a pretty core focus for us as a company, which is basically helping to unlock unstructured context to basically feed into these agents even as they get really, really good both in terms of the core model capabilities as well as the agent harness.
And we'll focus a lot on our core focus area, which is document context, but of course there's plenty of other sources of context as well. In fact, we do think context really is everything. In the end, you could have a infinitely smart agents, but your ability to actually get value out of this infinitely AI agent um, able to do more and more stuff end to end, is actually, you know, giving it the right things to do. Um, and the right things to do include the actual task, uh, the goal, um, but also access to the vast troves of organizational context available to it.
Whether that is web search, web context, whether that is connectors through, you know, like tools and skills and MCP servers, whether it's connectors through Snowflake or Databricks warehouses, um, or for us, whether it's access to the vast trove, like the 90% of documents that are stored within SharePoint, Box, Dropbox, S3, um, you know, all the unstructured, uh, data that's locked up within document containers. Um, and so for us, you know, we care a lot about how do you actually build the right tools and unlock the right context here.
We see documents as universal containers for unstructured context. Uh, there's over 10 trillion plus pages of human native knowledge locked up within PDFs, PowerPoints, Word documents, and Excel sheets. And at the same time, agents are also starting to generate exponentially more data in terms of, you know, more agent native formats in terms of markdown and HTML. For us, building an agent native document platform contains three main pieces.
Uh, it contains the document parsing layer, actually digitalizing, um, the vast troves of PDFs, PowerPoints, and Word documents into kind of accurate token efficient context, um, and feeding it, you know, either markdown, metadata, and other forms that actually enable agents to do stuff over your documents. The next is actually what we call the semantic and storage layer, um, which is kind of like this concept of document management for humans and agents.
Instead of humans opening up Microsoft Word, using your file system or file storage, how do you create some sort of agent interface that actually captures, stores all your documents, and enables agents to actually manage and operate over them in various ways. Um, and then there's still a room for actually kind of repeatable document workflows is agent layer they can actually encode within the uh within like this type of platform, too.
If there is a repeatable workflow, like invoice processing, KYC, or claims, instead of always offloading it to a generalized agent, how do you actually develop some sort of specialized workflow for it and really carefully tune cost and accuracy to make sure that you achieve your goal. So, when we talk about like the concepts today in in terms of these three layers, uh we'll talk about this concept of document OCR, which is one of the core concepts of this track, um to document extraction, document search, and document workflows.
There's still a variety of topics yet left unsolved from agent native document formats, document versioning, document editing, you know, hill climbing as a service, and just like a wide set of uh kind of remaining concepts to actually create a comprehensive piece of software where agents and humans can collaborate on documents and actually do work over them. So, we'll start with the first section um and maybe start kind of with uh the document OCR RP.
Um first off, uh you know, document OCR is hard. Um some of you might have seen a blog post that we put out a few months ago related to this topic. But, the reason this problem even exists is an agent cannot actually take the like the raw PDF like file binary and make sense of it. Um that's because the like the way PDFs are actually stored as a format is it's rendered for kind of like display purposes um and not really for uh you know, machine consumption in an interpretable manner.
Um they're designed for printing. So, basically text are represented as almost like individual glyphs with coordinates. Um tables are not represented as tables, they're represented as like typically line segments uh drawn with various types of borders, and also text drawn at certain cell positions. And so, if you're kind of an agent that's trying to make sense of this like document, you're going to have a really hard time actually trying to reason about what character and what shapes like map to what.
Um so the whole point of document OCR, um, you know, it's been around for like 20 plus years, is to really try to create some sort of, uh, digitalized well-interpretable, uh, representation um, that's both interpretable to humans, um, as well as AI agents. Also, reading order itself, if you have a multi-column layout, there is no, uh, guarantee that the way it's represented in the PDF, um, actually corresponds to the typical ways that humans would read it.
Cuz again, it's basically just an arbitrary sequence of characters drawn with uh, coordinate positions. Related to this, you know, even Word doc, uh, parsing is hard. Um, they're a little bit more structured than PDFs, but they're in kind of like this, uh, custom bespoke XML format, and this applies to PowerPoints as well. Um, it contains more structural information, but still there's a ton of like fluff. Like you don't actually need to ingest all the tags to actually have the agent make sense of the Word document.
It still needs to infer a bunch of structure from it. It needs like the right, uh, kind of to lift the right metadata around like formatting, semantics, that type of stuff. But also being able to ignore the tags and actually be able to render the document so that the agent can see the overall structure of the page. Um, it's still generally hard problem, and there's a lot of uplift you can get by actually parsing it into a more interpretable format instead of just using the native, uh, OXML and feeding that to an agent.
Um, if you're familiar with document understanding, um, you know, it's been around for quite a bit of time. Um, there's a lot of these like heuristic and pipeline-based approaches, which, uh, focused on kind of more, uh, I guess like human-driven, hand-handwritten like techniques to analyze like kind of various pieces of text, group them into clusters, and identify tables, paragraphs, and be able to kind of like, uh, generate some sort of output representation.
Um, a lot of basic techniques if you use open-source libraries like PyPDF, PyMuPDF, um, and of course like some of the, uh, more recent approaches, um, basically use this type of approach. There's also, of course, using a VLM to one-shot a document um, into text, uh, that's what we call like a vision-based approach. This works decently well. Obviously, you know, I think there's kind of a lot of uplift you can get by being able to read the visual structure, but there's a lot of sub-optimal pieces about it.
It can hallucinate on text-only pages. It costs a ton of money, and also it still lacks a lot of the semantics and grounding that you typically expect with some sort of document processing tool. And so for us, we really think about combining both the pipeline-based approaches of deeply understanding the file containers and binaries with the vision-based approaches to help generate, you know, kind of a hybrid approach that we think is at the Pareto frontier of cost and accuracy.
A quick note on this, I'll probably just like skim the high level, but the Pareto frontier for document OCR will always be like much more accurate and cheap compared to the Pareto frontier for, you know, wherever the frontier models are in terms of document understanding. It's because it's a very specific data type, and there's always ways to kind of like distill the latest visual understanding capabilities from, you know, Gemini, GPT, Opus into kind of a carefully tailored workflow that's able to process your documents at scale and in a highly accurate manner.
To some extent, that's exactly what we do. You know, we both optimize underlying like PDF engines plus like Word, PowerPoint, and others. We have an agentic harness that's like carefully tuned for auto routing between cheaper specialized models to frontier models, and also kind of specialized fine-tuned document VLMs that are parameter efficient and focus on specific classes of documents, elements like tables, charts, and others.
That's our commercial service called LlamaParse, which I'll talk about in a bit. But I think in general, you know, we're extremely committed to advancing the frontier of just document understanding cuz basically if you're within an enterprise organization and you have a massive long tail of documents across financial services, insurance, manufacturing, legal, government, and a bunch of others, there's just a lot of complexity in a lot of these document types.
And if you're actually trying to unlock context at scale, most of the models are not up for the task. And so, we've created this thing called ParseBench, which is a comprehensive enterprise document benchmark for agents. We we think about it as the most comprehensive enterprise document benchmark. You know, it contains 2,000 human-verified pages. It measures tables, charts, content faithfulness, semantic formatting. And it's optimized for just how AI agents are actually able to understand these documents instead of like syntactic correctness.
We've If you look at parsebench.ai, which is, you know, kind of the it's a fully public page on the internet. It's also available on Hugging Face and Kaggle. We benchmark probably like 50 different frontier models, open weight models, specialized OCR solutions. And you really want the Pareto curve to kind of be towards the left and up in terms of accuracy, like extremely high accuracy, but also extremely low cost. And there's just so many different types of documents, where some are a little bit simpler, maybe some are a little bit more complex.
And ideally, you want to cover all the points on the curve to deal with the dynamic distribution of various types of documents out there. It's pretty clear, even if you look at this graph, that it's like you can increase the complexity of the benchmark, and that document understanding is definitely not a 100% solved. But, you know, you fundamentally need to kind of advance a lot of the core capabilities to make sure they're able to process, unlock the vast trove of enterprise context out there.
There's kind of like a few different points on this accuracy cost latency Pareto curve. There's what I call like the high accuracy regime, where like, you know, some institutions basically need like 99 to almost 100% accuracy, because basically, the downside of an incorrect extraction is you completely mess up your financial model, you completely, you know, you basically get flagged for fraud or a bunch of other really really bad things.
In regulated industries like insurance and financial services, we see this a decent amount. This typically means you're willing to pay a little bit more money per page for like deeper agent tech reasoning to at least make sure that you get back the the information in the right format. There's also like the low cost regime. Let's say you're just trying to index, you know, the million plus documents per day within your that's you know, being continually updated within your SharePoint just for like rag knowledge base search.
You know, in these cases, you obviously want it to be not like terribly inaccurate, but even if it's a little bit messed up, it's okay, too. Because in the end if you have a sufficiently good agent, it can always dive deeper into the document and surface the right information with the right citations and grounding. So for these, you know, being able to create some sort of scalable offline indexing pipeline that has the best like cost constraints is something that is optimal.
One thing about VLMs based approaches though is that, you know, they're typically not very fast. And I think a lot of times if you have like real-time file uploads, let's say you're using Cloud Co-work, you upload a thousand documents and you need to process it within a minute a minute, like having a bunch of VLMs process that at scale is really tough for basically every single OCR service out there. And and to be fair, that includes ours, too.
I think in general, there's also some sort of need for an extremely low latency solution so that you can actually process stuff in real-time even if you have like deeper VLM enabled processing for kind of like deeper visual inspection and analysis. And so that's what I call kind of being in the agent loop. So besides Lama Parse, which is kind of our commercial service around like document processing and extraction, we also created this tool called Light Parse.
It is surprisingly really really good. I don't know if you've been following some of the Twitter threads, but it is Rust-based. It is the fastest open-source parser out there. It is completely free, um, and there is basically no strings attached. I think it's like MIT or Apache license. Um, and it basically is the most accurate like markdown parser out there that doesn't use a VLM or any sort of kind of like deeper model.
Um, and so this is kind of nice because you can use it as a default in the assistive agent loop. Um, let's say you're uploading a a bunch of documents to Quad Code, Quad Code Work, Codex, and you want it to process like a thousand PDFs extremely quickly. You can always do that, um, and then, you know, equip a VLM-based parser like Llama Parser or other frontier models as a tool. So, what these agents will do is they'll do like a fast pass over all the documents first, uh, uh, kind of like just scan through all the context extremely efficiently, um, and then if actually needs to dive into a page with like tables, with like charts, and actually needs to more deeply understand the values, um, it will use a VLM-based tool, slower processing, to actually make sure it reads the information correctly.
This is available as a one-click installable skill, um, complements kind of any other deeper VLM-based OCR tool you want to use, um, and we kind of designed to make it as fast as possible and also easy to plug in to your favorite AI agent. So, I kind of speed ran through a bunch of this stuff, but basically, you know, uh, we spent a bunch of time on the, uh, parsing layer. There's also other, um, general components around like the semantic and storage layer in terms of document extraction and search, and of course like document workflows.
Um, and due to time, I'll probably kind of just, uh, skip some of the, uh, unexplored areas like agent native document formats, hill climbing as a service, and others, but I'll share the full set of slides online. In terms of the semantic and storage layer, you know, besides document parsing, um, a lot of use cases also require actually getting back, uh, structured information at scale from documents. Whether you're processing, you know, a million invoices or expense reports or receipts or claims, you need to make sure that, you know, you want to actually get back structured outputs that you can put into a downstream database or system.
And so, a lot of these use cases basically revolve around the form of, you know, how do you automate a lot of workloads that humans typically do in scanning a lot of paperwork and doing data entry. Whether it is kind of, again, invoices, claims, contracts, receipts, or others. We kind of created these capabilities within Llama Parse as well. A lot of our capabilities are actually tuned towards like low cost while extremely high accuracy.
And you get back granular citations all the way back to the source document for every extracted output. And of course, you can run this in a pipeline at scale with confidence scores and also, you know, being able to actually flag whether or not we're certain about a certain value before deciding to put it into some sort of system. There's also document search, which is a basically expanded tool set as I mentioned around retrieval, BM25, grep, vector search, reading, and scrolling.
And so, all these capabilities are available within some of our commercial platforms as well as open source offerings. But I also just wanted to paint a picture of the general concepts out there today. So, I'll skip this section about kind of what's next and then maybe just go all the way to the end. I know I'm a little bit over time. So, really appreciate you all spending time today and then let me just how do I get to this part really quick?
Oh, right. I'm going to skip this piece. I'll put put this online. Our booth is at LG 47. If you guys are interested in stopping by and we're hosting a giant pickleball tournament today. So, thank you for your time. >> [applause]
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.