Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Doug Turnbull · @softwaredoug
Words
10,193
Runtime
1:07:29
Speaking pace
151wpm
Reading time
42min
151 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
All right. Well, why don't we get started? So, uh, for those who don't know, my name is Doug Turnbull. I am basically software Doug everywhere. Softwared.com is my blog. maven.com/softwaredoug. And this is a part of a series of talks I'm doing called retrieval augmented gathering. It's really a chance to uh discuss retrieval topics, talk about all kinds of things uh related to retrieval, rag, agentic search, and that's really what we're going to
76 words, the words spoken in the first 30 seconds at 151 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 510 |
| Average words per sentence | 20.0 |
| Longest sentence | 118 words |
| Questions asked | 52 |
| Sentences containing a number | 50 |
Most used terms
Filler phrases
496 in total: like 159 · uh 134 · um 89 · kind of 33 · actually 26 · right? 21 · sort of 18 · you know 8 · basically 5 · I mean 3.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
All right. Well, why don't we get started? So, uh, for those who don't know, my name is Doug Turnbull. I am basically software Doug everywhere. Softwared.com is my blog. maven.com/softwaredoug. And this is a part of a series of talks I'm doing called retrieval augmented gathering. It's really a chance to uh discuss retrieval topics, talk about all kinds of things uh related to retrieval, rag, agentic search, and that's really what we're going to focus on.
Um it's all in the leadup to a course I teach. I've taught probably a thousand plus people at different organizations how to improve retrieval and rag uh for the last several years. Um, I have a course on Maven called Cheat Search with Agents. I know some people here are alums of the course. Um, and love having you all back. But if you want to come and hang out with me and learn to do some great retrieval or how to build a gentic search in a way that actually helps your agents find things and not just BS you that they've found the right thing because we all know agents are amazing at gaslighting.
Um, definitely come check it out. And it's actually starting in a week. So, uh, the next cohort will be in a week. The probably be the last cohort of this year. Um, and [clears throat] there's a discount code here, $300 off with embeddings, uh, code embeddings. So, I'm going to talk about why embeddings are not enough for Rag. Uh, embeddings don't solve RAG. Uh that's today. Tomorrow we're going to hear from Daniel Tonka Lang who is going to talk about uh where he sees the information retrieval space going uh RA with basically his perspective on what he's seen with SIGR uh after going to the conference SIGR which is like the big academic information retrieval conference uh that happens once a year and yeah he's going to share his thoughts and perspective on where he sees agents agent gantic search rag and everything going.
For those of you who don't know, Daniel is sort of like Mr. Queryunderstanding queryunderstanding.com and just has phenomenal content on uh on retrieval and he's worked at search at places like LinkedIn, Google, Penda, those kinds of places. On Thursday, I'm going to talk about building agentic memory with Turbopuffer. So, that'll be a lot of fun. And uh yeah uh we had a talk from Adam Heavenor on Friday and there's a lot of similarities but I think I'm going to actually take things from a different perspective.
How do we actually build good retrieval and rag uh targeting helping an agent explore in a context and token efficient way basically your coding agents past history. uh but that can of course be extrapolated to um to any uh any um any I think agentic agent whoops agentic use case. So, uh uh the again the uh this I left the link off but of course and I will type it in here really quickly so everyone has it because some people are DMing me for it. mafood.com software Doug.
Yeah. So hopefully that works. Yep, that link works. And if you're interested in the course, come hang out with me. It's several weeks, three weeks. We have lots of guest talks on retrieval. And if this topic interests you, definitely check it out. So let's talk about the subject of today. Okay, I'm going to talk about why agents are not sufficient or why embeddings are a phenomenal similarity system but are just incomplete when it comes to actually building rag.
And there's two ways to think about this topic. One is critiquing embeddings just as a search technology. um their sort of limitations, their pros and cons, different limitations. That's one way to think about this. Another way is when we actually get out of just embedding like retrieval in general and focus on the problem of how agents actually search and then we can really see how embeddings are not a complete solution.
So, first I want to consider and just discuss what an embedding model is actually doing. And by the way, feel free. I'm very open to like being interrupted with questions. I'm very open to having, you know, raising hands. Um, a lot of this, of course, I'll cover in my course in greater depth. If this if you see something here you want to go deeper, that's probably a good path. But feel free to raise your hand and just chive it.
Um, if we think about the task of an embedding model, we have some documents or some chunks or some passages, pieces of text. I'm going to say text, but uh we could also be dealing with multimodal embeddings and have to deal with text plus images or just image embeddings, lots of things like that. But we know that there are there is some passage in this case something about BM25 the the search ranking algorithm and some queries what is BM25 and a embedding algorithm is a trained model that I've labeled here as encoder that after seeing many many many examples of passages uh and what's relevant and not relevant for and also queries on what passages are relevant or not relevant.
Those models are able to produce a vector, right? Uh and it should produce a vector so that they are similar. These two vectors are similar when the relevant passage for a relevant passage and a query. So what is BM25 should be relatively similar to this passage BM25 the book BM25 is an algorithm that blah blah blah blah blah and you can kind of eyeball these two vectors I'm just doing three dimensions of course embedding models are often 768 4,000 dimensions three dimensions is pretty easy to think about but you can just eyeball these and see oh they're kind of close together.
The each of these uh each dimension here is relatively close. The model has succeeded in producing a embedding for the query and an embedding for its rel relevant passage that are close to each other. And um if these if these vectors are normalized meaning their magnitude is one then we a simple we can just take a dotproduct which is just multiplying each dimension with each other and summing them and that gets what's called cosine similarity. we have a a similarity from -1 to 1.
And in this case, we see they are very similar because these two each dimension being similar is going to be additive to the final similarity calculation. And this is um what we expect how we expect embedding models to work. We expect before we index to embed a lot of our documents and then we put millions and billions of documents into a vector database like we elastic search turbo puffer whatever you're using quadrant and [clears throat] then at query time we do that same embedding with our query uh and we um get an embedding out and we are able to use that into the vector database to find the most relevant by some similarity measure most relevant uh passages and just by like laws of scaling like we have really really gotten good at this both in the vector database side and in the training of embedding models.
It's all though a function of our training data because these encoders these are encoders are billions of parameters that have to memorize not memorize or generalize patterns that it sees from training data to be able to output vectors that should be close to each other. And if our training data, for example, was entirely medical data, we might or in this case maybe entirely data about information retrieval, we might trust that it's going to do a pretty good job of getting a vector for uh content and a vector for queries in that domain that are close to each other.
If we take that same embedding model and then we use it on legal data, maybe we couldn't trust that. And so we're always making when we look at these embedding models, we're always making these trade-offs. What is this trained on? Is it trained on my domain? Uh could it be possible that it's such a large model that it could memorize many domains? Um and uh is that is it appropriate for our domain? So that's always the question we want to ask ourselves when working with an embedding model.
Is it like an appropriate embedding model for our domain for the kinds of problems we're working on or is it specialized? Uh these are things that come up all the time. For example, maybe whatever embedding model doesn't work well in e-commerce, we put in a living room couch as in at and try to encode that document, a blue living room couch. And um we for the query when we encode that into a vector they just produce vectors that aren't actually similar.
And you can kind of eyeball these and see they're not actually similar. So it can happen that if you choose a poor embedding model for your domain, it's not going to potentially rank things well for your domain. A way of thinking about this is a kind of compression. So, if you if you look on the left, we obviously see a blue couch. But it's very easy to get lost in all the granular details. If you were to compare this blue couch with this blue couch, you might say uh the ones on the left, you might say they're they're fairly dissim they're fairly dissimilar from each other.
They have very different couch attributes. when you kind of fuzzify them and compress them, which is a way embedding is taking all this text and kind of compressing it into a a single vector, you just kind of get if you just squint your eyes when you look at you're like, "Oh, that's a blue couch and that's a blue couch." You've kind of reduced that's sometimes people use the term reducing dimensionality down to the core attributes that actually matter to our training data.
Um, and depending on, you know, I mentioned before an embedding model could be 128 dimensions, it could be 4,000 dimensions. When we add more dimensions to our embedding model, notice now, let's say this is a slightly less fuzzy couch image, we can start to see more of the features of the couch compared to this one that takes away a lot of the features. So, um, we start to see more like maybe if we searched for tufted, uh, tufted are these little like dots as far as I understand or rounded edges, these like rounded corners here.
Um, we might actually be able to recover some of that information. So there's value the basic idea a lot of the basic idea of an embedding there's value in compressing information but kind of compressing it towards the the types of queries the way that the things that we're going to use to measure similarity. So if our queries are about toughs or rounded edges then maybe we are our our fuzziness kind of goes in a specific direction and not like a different kind of direction.
So we're kind of emphasizing specific features when we're compressing. So like I said, this may or may not be uh if you took a state-of-the-art embedding model. So a very common state-of-the-art embedding model is Voyage 4 large. It's a very common model or I would say not state-of-the-art but pretty good is one of the open AI models and you just like used it on your text. You have to ask yourself does this have is the query here the query document relationship well a is your problem going to be in the training set even if e-commerce data let's say you have an e-commerce problem even if e-commerce data is in the training set there are many types of e-commerce problems like groceries work differently than uh selling books um and not only that even within a like a bookstore, what your users expect when they type something in might be slightly different than the general expectation.
What if blue doesn't mean color? Like purple is a brand of mattress, for example, in the US. Uh at Reddit, for example, we had all kinds of challenges with like slang that would come up suddenly or news topics that would change constantly. So, uh, the state-of-the-art embedding model was not necessarily going to help us because that language was evolving and how people searched evolved all the time. Um, like I said before, is the model large enough to get all is the model large enough?
Is the dimensionality large enough to get all of the features that are important to your domain? Like all those little parts of the couch that we showed. Um, and are we like able to encode the PI are we able to get enough resolution of that picture to get like the fuzziness we need between query and document. That all said, it's very hard to beat if you take Voyage for large. These are all notebooks which I'll send out uh in my email after the class.
These are all using turbopuffer. Uh voyage for large. Turbopuffer added a great feature to uh just let you embed. You don't have to call voyage separately. They just do the embedding for you. If you embed voyage for large and you compare it to BM25, it's almost all across the board. The voyage for large on average will beat just uh BM25 search. But the devil is in the details. So about if you look at Wayfair Wands, Wayfair Wands is a data set that is about furniture shopping. you see a lot of queries that are very badly harmed about if you there's about uh I would say I can't remember the exact numbers but like 30 roughly if you compare the voyage 4 large to BM25 roughly 33 queries were harmed and 66 got better with voyage 4 so it's not like a slam dunk and one thing as a bit of an aside I'm always hesitant to ship changes to production when the there's such a like there is a significant portion of harmed queries.
That's probably a sign that we need to do a better job of routing the right queries to the solution that actually might benefit from an embedding model. So all to say even with some open data sets it's not a slam dunk that every query is going to be improved. On average you'll see improvements but in some data sets you'll see like what's interesting with voyage for large some data sets that are very information dense like scfact and um trecoid those are two data sets that are very focused on tcoid is like questions about covid and then that came up during covid and then scfact is like scientific um scientific uh fact debunking of facts, that kind of thing.
Those uh it's very it's a clearer win for Voyage 4, but for some of these other kinds of query use cases, it's much less so. Um [snorts] and then we could ask ourselves, do we the next steps, would we find a more appropriate model that specific to this domain? That's our what we need to think about. Or should we implement some kind of hybrid search? we might just want to do BM25 too. Like I was quoting BM25 before, but just to define a little bit why it's an attractive baseline option.
Um BM25 is nice because it's effectively doing a kind of per token. It's a keyword search. So every we're just summing all of the occurrences of of the terms that match. And every term that matches has a weight that's that's computed based on what BM25 assumes the importance of that token is to the user searching. So uh that's all done via statistics that are ultimately based on an assumption that for example rare terms in the corpus if a user types them out and they're searching with them are more important to the user's intent than very common terms.
So baby name might baby might be more rarer than what obviously like what is such a generic term. baby or baby names is only going to occur in very specific contexts. And because of that, because of that specificity that uh BM25 brings, um it tends to be a very defensible ranking that doesn't need to be trained on anything uh and doesn't depend on whatever your embedding model happens to be trained on. So, it's often a great idea for things that are outside of open data sets or your specific problem to be able to uh also implement BM25 as a potential solution.
Um, so I'll stop there for questions. I know I just went through a lot. Feel free to raise your hands. I think what I would say is in summation BM like embedding models depend a lot they're compressing the space down but they're going to depend a lot on what their training set is and the quality of that training set increasingly those are ext like the voyage for training and like how what it's able to do is incredible but even there you're trying to represent every domain's meaning in a24 dimens vector and sometimes it's just going to be off a little for one because it has to make trade-offs to do that compression.
How it's going to fuzzify things is going to be different that it's going to hurt some domains and help others. So yeah, any questions I could help answer? >> Um so I have a question. So for uh for for uh >> um I mean for a general language corpus right pearl or let's say which is um let's say a data set which is not very specific for example in this case about furniture right are there any um any any embedding models which are rated like really high in terms of gen general language understanding you did mention voyage for large in in this case but are there any ones which you personally um prefer or are um or or could recommend? >> Yeah, I mean for general language for like generalformational search where the goal is to just do I would say question answering kinds of things like it's hard to beat.
The voyage models are hard to beat. I think they're also very expensive. They're more expensive. They're not like very expensive. Um but uh what I'm finding more and more so there's like there's this trade-off, right? There are these general models like Voyage or the OpenAI models that are going to have a massive amount. They're going to have billions of parameters. No one could ever run those in their own infrastructure or it would be very hard.
Um and those are going to be for general language just um really good. And then you have I would say smaller or more targeted models like the ones you find on hugging face often which like um an embedding model that's just focused on medical data or an embedding model like I used to use at a past company that's just focused on fashion images and those are more specialized and you can self-host those more easily. Um, and it's sort of a a trade-off of like do you take something that's generally good and going to be good at general information or do you need something that's like really specific to your domain?
Uh, and there's no real good way to evaluate that other than to get some data and do some evals yourself. >> Okay, thank you. Um AJ has interesting question. We are mapping normalized job titles into a vector space using open AI taxing three. The model conflates seniority with functional domain assigning high vector similar to VP of sales and CRO due to executive tokens index near index zero. Um that is going to happen for sure.
We tried schemadriven augmentation by extracting structural attributes role level function domain by but flattening those fields into a single text string before embedding didn't stop representation collapse. Yeah, the self attention layer still overindex on seniority tokens. Yeah, this can h so um what AJ is talking about can happen a lot where it's hard to anticipate how embeddings prioritize different things and not other things.
And I'm going to talk about this in a little bit. Um but a lot of these structural types of um of search and matching problems you almost need I I think we we run into problems. This is actually an area where embeddings might not be the best solution. Uh, and there's a completely different solution to doing semantic search based on hierarchical classification that requires you to build a taxonomy. Like if you build a taxonomy of titles for a specific domain, and you can't do this for every domain, um, that can also be like an underrated way of doing semantic search.
Uh that's kind of a a tangent. Maybe I'll talk about that another time. But um uh you need when you're doing your structured attributes, it's not just about um not just about the embedding per se and like trying to hack around the embedding, but as we'll talk about in a little bit, it's also about being able to filter the embeddings and be able to have a notional semantic way of doing filtering at different levels of specificity of like job title, if [snorts] that makes sense.
So I would how I would restructure your architecture is probably keep the embeddings as raw similarity but if there's something where I need to um uh prioritize and I would need to classify content. So I built a classifier into a large I don't know if you can model every job title that you care about. Uh but there's like extreme multilel classification is sort of another end of the spectrum when it comes to semantic search if you want something to search for.
[clears throat] >> Thank you. Do >> Yeah, sure. Cool. So, uh another thing that happens a lot uh embedding collapse hubess there's a lot of names for this. If you take the other the one big downside of a generic embedding model and this actually also gets to AJ's question. If you just take a generic embedding model and you embed your domain, your domain like if our domain is furniture, everything in our domain suddenly is close to each other.
So you kind of get shocked that everything like the similarity between a blue couch and a table is much closer to one than you might expect. Like you expect these two couches to be similar, but uh these up here, the table and the couch are not. And this is another consequence of it not being specialized to your domain. It's a general embedding model has to know lots about everything. And so it might not it may not differentiate between really nitty-gritty features in your domain.
It also may uh just sort of the distances between those things gets really small and surprisingly close. So you have some kind of like for example sane cutoff that you add um where you're trying to have like something that's ranked like this. Again, here's another thing where it's like, do what's more important? Getting the couch right in the similarity or getting the color of the couch of the item, right? Um, and a lot of people will spend a lot of time trying to find the cutoff for embedding models and what what embedding model similarity cutoff should be.
How do you draw the line? How do you find that threshold? And I've been in situations where I've like tried to learn for a given embedding model and a given type of query what the right threshold is to like filter things out of the results and include other things. And man, that is a that is a very difficult task because you have if you search for something generic things generally will not be super similar to generic terms of like couch.
But if you search for something really specific like blue couch with rounded edges and button seams blah blah blah, it's actually you will get like one or two things that are extremely similar and then like this massive drop off. So there's no clean like threshold. And one of the biggest flaws of embeddings, I'm not sure it's a flaw, it's just like a limitation that people don't necessarily appreciate. Embeddings don't do set math.
So a very common thing that could happen and this happens to me all the time when I work on search um client comes to me and shows me this this query's results they say I search for blue couch and I got a lot of things that were not blue couches even if the top result was a blue couch and it's ranking perfectly being they want to not see things that are not blue couches and that is a common thing across the board that I've seen in many domains where people are surprised that this becomes searches behavior.
But it's also one of those things that search people are often um in a position where they're like uh they think their job is to get ranking right. And our job isn't just to get ranking right. It's also to select the set of things to show the user and the set of things to not show the user. So that's right uh Brashan. There's no there's no easy threshold for embeddings. I often wonder and someone could correct me or tell me if there are embedding models that try to build this kind of threshold cuz I bet if you change the training task, you might be able to have something like that.
But it's important that embeddings are similarity systems. They're not classification systems. And so uh this is uh yeah that's so suja brings up a really good point of having a you can actually do you can build a classifier out of embeddings uh but it's it's going to be like calibrating and it's also going to be very query dependent in my experience. Um, Suji, you should write an article about that and then I'll ask you to come to a talk. >> Yeah.
Well, yeah, this was like a data set calibration thing, but I found it successful. So, yeah, >> cool. >> I could maybe. Yeah. Uh so it's often the case and this is very this is something that's going to matter a lot for some domains a little for a lot of domains and there may be domains where it doesn't make any sense but uh if you do care about only including things that are relevant excluding things from at all that are irrelevant query and content understanding still matters a lot and of course we've had in the last uh week or two, Jev has come out and I think that's really exciting because we have now um we've now had this sort of explosion of realization that many of us have known for a while when we've used LLMs.
We've used them a lot for classification. Jev has just really fine-tuned towards that task and this these system one models Jev and all the different open Jev clones. So, I suspect we're going to see even more people aware that they can add structure to uh to their products and the things that they're searching. So, one thing to keep in mind is that um and there's a lot of great stuff on there. There's someone who uh on Twitter was talking about building an embedding.
Whoops. Building like a actual not blackbox embedding where every dimension was a question it would ask Jev. So is is this color blue? Could be a dimension in an embedding. Is the category couch could be a dimension in an embedding for example. And what Jeb will do is it will give you a probability that it's true or false. And that can just be another dimension in embedding. And what we're rediscovering is something that we've known for a long time is that you could have an arbitrary list of features that could be an embedding.
But in any case, uh the main thing is we should not discount the the work that is doing good content understanding and query understanding for the domains where it's really important. And often this is what users care about when they're searching. They want to know that you understand what blue is and what blue is not. They want you to understand what a sofa is versus uh versus these other things. So that's actually really important for users.
So the other thing about embeddings I like to think sometime embeddings are things especially when you're taking an off-the-shelf embedding. Embeddings are these like giant similarity moon lasers that we trained uh in some giant compute if we have the ability like we may have the ability to train embeddings in our workplaces but for the most part a lot of us are just taking an embedding off the shelf and we're reusing it.
Um, and that's something that uh lets us sort of just put some data in, get put in a query, and see some content come back that's similar. A thing that we often don't take for granted if you come from a more full text search background is how tunable a a search system is. So, in this case, this is a turbuffer query. And I found I could just spend a little bit of time dorking around with BM25 boosts uh and weighing things differently and just seeing what would happen if I played around with the wands data set where out of the box the embedding sort of won.
We got a pretty good score for the Voyage 4 data set. It seemed to beat BM25. And honestly, one great thing about working with a search engine is you can just dork around with how things are weighed or ranked. and recover a lot of that difference pretty quickly. So, another big downside to embedding is they're very frozen in amber when it just comes to with thinking about retrieval, they're frozen and amber perspectives on um on retrieval.
And what's great about just working with some of our traditional unsupervised basic lexical tuning systems and filters and query understanding all that stuff can just be liveetuned. It's almost like it's lazy lazy tuned relevance whereas uh an embedding it's like all encoded up front. So it's useful to know about these other tools. Um, let's see. I see Patty has a question. Uh, and then we'll move on to the some more fun stuff.
I work on a higher ed library as a metadata library and our search engine will move to an AI assistant. We're interested in hybrid models that have both traditional index search and embedding search. And what this means for me as a metadata librarian, um, my main concern right now is how to handle ranking results in good catalog records getting pushed down the page from non-catalog records. Would that mean we have to move away from embedding models and back to traditional search?
Uh, that could totally that could totally be the case. I'm not sure what the difference catal good catalog records being pushed down the page from non-catalog records. So, um I think what you're getting at is the idea that when we're doing embedding search, that kind of goes back to uh what we talked about before. There are many things that dictate whether or not something is relevant. I'm actually going to talk about this in a little bit when we talk about agents.
A lot of what what embedding models do is they're really amazing similarity systems. But a very important thing that I think is to take away from this talk is that similarity is not the same thing as relevance. So things can be very similar to a piece of text but not necessarily the right answer that you would want to show a user. That's not necessarily something where it's just about VM25 versus embeddings but it's something to think about. uh like print books have catalog records but don't always include an abstract and digital journal articles won't have a catalog records but have abstracts in their research record.
Yeah, in those situations uh if you have dis if you have incomplete data that's going to be a major problem regardless of I would expect that to be a major problem regardless of if you're searching abstracts with keyword search or embedding search. So that's just something to keep in mind is you may like a lot of people are using LMS for example to generate and fill in bits of data when they don't have that. So that's something else to think about.
Cool. Well, I want to talk about rag specifically. Uh so far I've kind of critiqued and shown you the limitations of embeddings as an overall ranking and retrieval service. But let's talk about RAD because I think this is really where the problems begin. I think a lot of people, especially if you work in retrieval, you'll quickly run into the limitations I talked about. You'll uh see that embeddings aren't a one one-stop fit for one size fit for retrieval and it's a more complicated solution than that.
But rag and agents are actually a bit more of a specific thing. So what we've all long been taught to do with rag is to chunk up like this book. Let's say we have a book of 100 pages on BM25 and we chunk it up into these chunks somehow. And they're they could overlap. They could have some kind of um they could overlap. They might just be based on length. They could be like 10 24 characters. They could be lots of things that we do.
And we embed them, right? So, we embed at this level and that helps us get at the the similarity or the the the meaning of this specific chunk because if we try to embed the entire book, that's not particularly useful to embed an entire book because a it's all of the uh passages. It's kind of like all the embeddings are going to average out to uh b it's going to be probably too large to embed at all fit into an embedding models context. and C um an LLM isn't going to be able to consume read and you know consume this entire book.
So we embed at the we chunk and then we embed these individual chunks and that's the idea so that we can retrieve and look at these individual chunks and not try to retrieve an entire book and shove it into the context and then fill up the context because the context window is only a million tokens wide. So we do what we did before. What is BM25? And we have one of these chunks that seems to match that, right? Um and then the main thing, the other thing we have to consider if we the chunking itself, if we have big chunks, that's going to dramatically change how the embedding that's going to come out of this chunk.
So if we start with a small chunk here, if this was the sentence, that's going to be very similar to this query. What is BM25? When we start adding information of where this came from, look at uh here's the answer, but look at the rest of the text. It's a lot of stuff that talks about history of where BM25 came from. If you didn't know, this is all like about the history of the system that it came from, where it was developed, blah blah blah.
So, if we add more of this context, in some ways, that's useful for us to understand uh where this sentence came from. On the other hand, it begins to nudge this initial sentence that was useful. Now we have this larger chunk and we nudge it away from what is BM25 and we increasingly are nudging it towards tell me about the history of BM25. So it's almost like a better fit for this query than it is for this query. So it creates a kind of weird problem where if we chunked a certain way we would be really good at answering one question and if we chunked a different way we would be good at answering a different question.
So what I tend to see is people pick these chunking strategies that tend to fit whatever their eval query set is. But uh it's often very hard to do this because adding or removing text just tends to like you're just always right sizing to your eval set and some things are going to be very targeted towards one one chunking strategy is going to work well for one query and another that same piece of text chunked differently is going to work better for a different query.
So I think people are often arguing about chunking strategies because and sometimes it's useful to have this bit of outer context and sometimes it's good to just focus on the individual sentence that uh that like is a disconnected fact in this corpus. Yeah. In one sense, we've diluted on the right hand side. We've sort of diluted the meaning of the original first sentence and the smaller one keeps the history keeps out the history context and just answers that question directly.
So there is this problem and I'll just stop here because I know I've gone for a while. Does anyone have any questions about what I'm going on? Does this make sense? Sort of how how chunking is going to dictate our ability to answer specific questions when it comes to retrieval. >> Um, so I have a question here, Doug. >> Yeah. >> Um, so in in a RAD system, right, assuming both of these sentences were were together were a part of one chunk.
And uh firstly, let's say uh if my query was what is BM25, wouldn't my my similarity search still be like able to pick up this this chunk and and uh like that's first part of the question because let's say if it is able to pick up this this chunk irrespectively eventually it is anyway going to hit the LLM for a generative response and the LLMs are pretty good at figuring ing at just telling you about even though you gave them the entire chunk, you should still be able to get that the answer is BM25 is a lexical similarity measure, right?
So, so uh I mean going back to the first part of the question um even if these two sentences were one chunk um is there a uh is is there a case where I I might not get it as an answer for the query what is BM25 like would my s um embedding model not be able to pull this chunk out in that case? Yeah, I think it would be I think a more interesting example than maybe even the one I have here is if we had a similar sentence defining BM25, but the rest of the text went on to keep defining it.
Um, uh, that would be a case where the rest of the context would push us towards this being a better answer to what is BM25 than the than the one that answered it quickly and then started talking about history. if that makes sense. So, uh I think you're right. It probably is going to be it's probably fine, but I think the point I'm making is just adding it's like we're adding a bit of noise >> that tends to move a chunk in one like depending how big the if we make the chunk a little bit bigger, we move it towards being able to answer one query a bit better.
And if it's a bit smaller, maybe it answers a more targeted query a bit better. And it's hard to know ahead of time which queries our system will get in any and it's probably going to get both. Right? So it's we're always trading off between almost this like inner I I like to think about it as and I'll talk about this in a in a second. We're trading off between this like inner fact and the outer context of where it came from when we think about chunking.
And if we by changing making our chunk size bigger, we're like adding more outer context that might dilute it or it might be good. It might we might be like steering towards um a better solution. It might be steering it towards like it actually being embedded better for being more what that where it came from and what it's about. >> Got it. So I I guess what what we're trying the point I believe what you're trying to um make is that we do not really have a handle on on how on what what goes into a chunk uh and we may not and and that that content may not be geared towards one query as much as towards another query and it may fit both or or Okay.
Okay. Got you. Thank you. >> Yeah. Yeah. Totally. Um, lots of good questions. I'm going to I know we got about 11 minutes left, so let me keep going. [snorts] So, one way people have solved this is through a technique called late chunking. And I'm going to talk about late chunking now. So, if you know how a language model works, so you take this piece of text and you turn it through a language model, every token in a language model has a vector.
It has a state. And that state can be used to answer questions like what token would ideally what would the embedding of the token that would usually go here? What is like how could I in a mass language model you can say like if I took you compute an embedding from this and you could say like something like BM25 should go in this position based on all of the words that are around it. And you can do this on like the full let's say we have a book chapter that this chunk is coming from.
You can do this on the full book chapter and get a state for every position in that book chapter that captures the meaning and the context of everything around it uh in that chapter. So in the example before uh I was talking about history, it's going to be BM25 but skewed a bit towards history. Um what the idea behind light chunking is if we did this if we took the before we did any chunking if we took the outer text and we ran it through a language model to get the uh that state that token state at each position we could use it to construct a better embedding.
So later when we do chunk it instead of chunking and sending it to a a vector model the idea is with late chunking we we chunk based on whatever and then we have these little artifacts. Remember these are like predictions that are based on the the full non-chunked version of this text. We can average this. It's called mean pooling. and that can be our vector for this chunk. So it's a completely different way and approach to thinking about embeddings and chunking.
Uh but it's a way of saying like preserving the outer context but also keeping uh but focusing in on like what actually was said in that inner piece of uh text. So late chunking is one thing that people do. It's still not foolproof. We're not necessarily, like I said, we're not pushing this through embedding. We're just like using the actual token states from a language model. And we're still we still have to make this we're still making trade-offs between inner context the thing that just directly answers the question and the outer context the broader picture of like this is in some history article not necessarily trying to teach you about BM25 and what it's good for.
Okay. The other approach that tries to capture try to trade off between inner context and outer context is um is something similar to uh Shashank is asking about coar but is is is this idea of late interaction. So lay interaction is the same idea. We have we take a document which in this case could already be chunked. It could be the larger larger piece or whatever. And um we put it through a late interaction model. And what a late interaction model is actually a bit different than a normal.
So in a similar sense, it's similar in the sense that every token has a an embedding kind of a a vector that's associated with that token, but the purpose of that token is different. So the purpose of the token when we're doing late interaction is to say okay uh I am going to emit tokens so that when a query here's a so here's a document BM25 is a lexical formula here is a query what is BM25 and now instead of in the traditional embedding sense we're doing query to document just one vector search between query and document Now every query that we're searching with has its own embedding and we do something with find what is BM25.
We find the token that is most similar here for each of these here. I'm just focused on BM25 and we see maybe it's actually this lexical one for whatever reason. Maybe the surrounding context is kind of saying this should be this should attract uh a BM25 asked in this context. So then we score based on doing this finding the best most similar uh query doc token or document token for every query token and summing that and I know like if you haven't messed with this stuff before it's if you haven't thought about lay interaction it can be like a big mind blowing experience but the main thing is whereas normal normal language models that that vector at that embedding is about trying to predict what token would go there.
I think you could think about the uh in the co bear situation you're not just trying to predict what token would go there. You're trying to sometimes more predict what query token might go for this document token. So I'll stop there and just both of these I know are just different ways of having uh a sort of preserving that outer context but in the inner unit of the chunk itself. Whoops. Uh Deshan asked when we do late interaction do we get keyword search implicitly as each token is now compared so the search term in the query will get maxim with the document tokens matching the same exact term well it's not BM notice how BM25 in this position like if we have BM25 mentioned multiple times in this piece of text what that's actually doing um they're going to have each of these is a unique what's called hidden state for each position.
It's not unique to each token. So, it's capturing something about the information that surrounds it. So, if you had what is BM25, BM25 or like BM25 is a lexical formula period. BM25 stinks. those two different mentions of BM25 are going to have different hidden states. They're going to have different vectors. So, uh not necessarily and you may have a situation where what really seems to tie it all together, for example, is this maybe lexical term because that's the meaning based on the all this is trained on training data. the meaning of flexical and what's predicted to go there seems to somehow go with the BM25 on the question side.
Hopefully that makes sense, but it's a it's a big topic. Don't feel bad if you don't if you're not like fully understanding late interaction or late chunking. main thing I want to get across is there are these techniques out there that people are thinking about because they realize that depending on chunking is going to change what we're retrieving and everything is very chunking dependent. So they're trying to find a lot of the retrieval community is trying to find ways of sort of not necessarily having to throw away this outer context when they're zeroing when they have this unique like paragraph or chunk from some inner context. >> [snorts] >> Um so the other thing that happens is we have to think about chunking.
A big mistake when we when it comes to rag is thinking about chunking only in terms of retrieval. Chunking is also how agents use helps agents use the information. So if we have this what is BM25 and we have these two passages at Shopify we worked on ranking one feature of BM25 BM25 little similarity blah blah blah now chunking is not just about something upstream that a retriever is caring about it's as much about how the LLM interprets the context uh because think about the an LLM is just going to get this and get these random paragraphs out of nowhere And then it has to answer the question, what is BM25?
And how does an agent decide what in this is useful? What in this should be thrown away? Like think about where the agent in this position is now a researcher. It needs to evaluate evidence and it needs to think about uh think about the sort of like pros and cons of each of these different passages. So, for example, if we had this raw chunk, BM25 is what we've been talking about and then it turned out it was written by George RR Martin in a software engineering book about Westeros or something weird.
If you're an agent or you're any kind of researcher, you're going to disregard this as being not particularly useful to a serious information retrieval researcher, right? Um, it has really nothing to do with that. So a lot of what we're trying to actually do with chunking is give the agent a representation of the information for it to make a good research decision. So instead of thinking about and I think in some ways we need to think of the agent in this context as almost like a reranker like the agent's job is to get a selection of a sample of the corpus and not just random disconnected chunks but like a lot of metadata about what makes those chunks interesting like why they might have been retrieved blah blah blah.
Um, and if we have all this stuff like publication date, who the author was, like its popularity, we the agent can then make a better decision about how to use this information. So the outer the getting that outer I I keep talking about this trade-off between inner and outer context. That's also not just about retrieval. It's also about giving agents information. So that outer context isn't just a little bit more of the paragraph that goes with a chunk.
It's also that like a section or a chapter heading. It's also like where did this come from and all of this other information. So agents need to not un not just see retrieval fetch something. It often needs to see why retrieval fetch something. >> So then it can decide on relevance. Yeah, I think this is a copy paste from before. Um, and search is like I said, search is really about agents ability to evaluate chunks.
And really what we're trying to do when we're doing chunking, um, and I think the emphasis on chunking on like figuring out where to cut off some text is is one of the biggest problems that people that that happens in rag. People think it's all about retrieval, but really a lot of what we need to be doing with chunking and why to tie it back to the original part of the talk, why embeddings aren't enough for rag and why this sort of chunk plus embeddings regime isn't enough for rag.
We really need to what our real job is to help agents see forests and trees. So the tree is that specific piece of context and the forest is like where it comes from and how we might get around the forest. And also when we are searching the forest, when we're just like striking out for a general topic, we need to be able to give a broad sampling of the forest and not just like accidentally give like 50 results about that written by George R.
Martin on the software engineering of Westerosps or something weird. So diversity is extremely important when the when the agent is not being very specific about what it wants. So I'm just going to tie this up because I know we're at time on like how I tend to think about this stuff. Uh first I really think it's important to think about as much per keeping as much per domain structure and using chunkings as almost like a last resort to satisfy like some kind of constraint on embeddings.
So I'd rather see a large book broken up for set of chapters and then sections than just like a large book going straight to chunks. Chunks become the sort of last resort. And then when we do that, we can give agents a better sense of the structure of what came back. Think about one thing I haven't mentioned. Think about like six months ago to a year ago, everyone was talking about how GP was all search was going to be written in GP from now on.
I feel like that hype train has kind of come and gone. Um the reason for that is because we often organize documents in the same way on a file system in a way that made sense to us as humans uh and agents were able to maintain that organization to some extent. Same similar thing here with breaking things up into chapters and sections or whatever. Uh but we show that structure to the agent when we're returning results.
And then when agents are aware of the structure of the content, we give them tools. We um that's in some ways more of a navigational thing. First, we searched for a given um title or we searched for a given keyword and the uh we got some diverse sets of results back and then eventually we were able to help the agent sort of gets filter down to whatever book it thinks it's important to search into. So in that way agents are sort of thinking about the search process as a stateful thing that they're navigating through.
And then of course diversifying when getting general searches. And you see this all the time even in agents when they use web search. They will strike out with a broad query and then they will get very specific just using keywords even uh getting into nitty-gritty details of like first I want to go in this direction then I want to go in that direction. Um and in that way uh you start by getting this picture of the forest and then give agents the ability to get deeper into trees.
And I'll leave it there but in my perspective like rethinking search away from this thing that's just a flat set of chunks that are embedded is very important and a big theme of my cheat at search course which I hope everyone comes to in one week. So, I'll stop there. I know there's been a lot of great chatter in the uh in chat. Uh and feel free, do people want to raise their hands? I'm not sure what questions are still current.
I'm happy to hang around. >> Yeah, Suja. So, so one thing I was uh thinking like to your point about uh features, right, that really when you compress it down to an embedding model, you don't know what each element is, right? Because you don't know how it compressed. One idea that I was thinking of while you were talking is to take the lexical sparse vector maybe like a 5 million long vector, right? Uh which is sparse and then doing a variant of uh you know the old uh what's that called? um the the Stanford thing, right?
The word tove versus the other one, right? >> Um doing that like a PCA on it, right? And bringing it down to maybe like a you know a 40,000 30,000 split size, right? >> And then what you have left are the ones >> the you know the heavyhitting uh elements, right? The >> they're still terms but they're the important terms in your corpus, right? M >> and then you could have explainable um kind of vectors. Um >> Oh, interesting. >> Yeah.
And also semantic because you have basically squished the thing, right? You have done that vector multiplication >> sort of like halfway between sparse and dense vectors. >> Yes. Yes. Yes. So that's something I was just thinking of that maybe >> that's really interesting. You almost have like you take a vocabulary and you compress it, but you don't compress it to the point of it being like so small. But even but you can still explain each dimension as almost associated with a specific kind of word or theme, >> right?
Well, you don't even compress it. You just rotate the axis, right? So that each of sparse elements now become dense but very small number of elements, >> right? Those are the so-called important ones. >> Uh-huh. Yeah, that's interesting. >> Well, I haven't tried it, but you know, >> you should try. Okay, >> I'm going to get you to come back and present that. I think that's a good idea, though. >> Yeah. Another thing I didn't uh you didn't cover, I think, and I think it's very interesting >> is when uh you took features and then this is when you are assigning weights to BM25.
And maybe this is out of scope for this article, >> but you use uh some kind of logistic regression to basically figure out the weights and do a like exhaustive work or a random work on the search space, >> right? Remember with the weights. So I I think you blogged about it and I thought that >> Oh, I have done that like a like a basian search. Yes. On the weights. >> Yeah. >> I haven't I didn't do that here. Like I just >> tried something simpler.
Yeah, >> I think it's a great idea like well honestly increasingly these days so if I really want to get exact I'll do a basian search but like I also think so what Sue just asking about is like for weights like this this one in passage or title you can optimize those like you can just try different ones and see which one gives you the best like ndcg um honestly increasingly I just try a random search and it works pretty well >> for simple things So I'll just like randomly try zero to different values 0 to 10 until I get to something I like.
And I've been pretty happy with that for a lot of things. Uh and it may be that I I don't know it may be that the basian part might not make sense. So you have a lot of parameters. Any other questions? Cool. I know we're at time. I know there was good chatter. Cool. Well, I will uh please come to my course if this kind of stuff interests you and you want to get hands-on with building a gentic search rag using LLMs and Jev and whatnot to do some fun stuff.
So, all right. I hope everyone has a good rest of your day. Take care, everyone. >> Cool. Thank you. Yeah. >> Cheers. Yeah.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.