Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
Most replayed moment at 13:49
2.3x that video's typical replay level
Fable uh and it runs into an unknown, ask it to log it, right? So that um you uh you can see where the deviations happened and then you can sort of figure out why as well, you know? It will usually give you some context about what happened.
Said at 13:43
The graph counts replays. It does not show where viewers stopped watching.
Words
4,024
Runtime
20:11
Speaking pace
199wpm
Reading time
17min
199 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
I'm on time. Hi everyone. I am Shivam. I'm from Spotify and my talk is going to be about how at Spotify we do personalization especially in the era of LLM's. So uh for those of you who use Spotify I guess do we have any Spotify users in in the room? Raise your hand. Nice. Uh yeah, thanks for using Spotify and a bit like in this talk so this is going to be less about context engineering from the conventional like agentic sense. It's going to be more about how we do context engineering on the modeling side.
100 words, the words spoken in the first 30 seconds at 199 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 179 |
| Average words per sentence | 22.5 |
| Longest sentence | 80 words |
| Questions asked | 6 |
| Sentences containing a number | 6 |
Most used terms
Filler phrases
416 in total: like 162 · uh 83 · um 64 · kind of 63 · sort of 32 · actually 7 · you know 4 · right? 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
I'm on time. Hi everyone. I am Shivam. I'm from Spotify and my talk is going to be about how at Spotify we do personalization especially in the era of LLM's. So uh for those of you who use Spotify I guess do we have any Spotify users in in the room? Raise your hand. Nice. Uh yeah, thanks for using Spotify and a bit like in this talk so this is going to be less about context engineering from the conventional like agentic sense.
It's going to be more about how we do context engineering on the modeling side. So if you're interested at all on how your Spotify app works, how we recommend you songs, tracks, episodes, etc. This talk is going to be really useful for you to contextualize how we use your data for for just recommending stuff to you that you like. Um so a bit about me I am the tech lead of the user representations team in Spotify's AI foundation org.
So the AI foundation team builds all of the frontier foundational models that that are used across the entire stack of recommendations at Spotify. So we do things like user representations, content representations as well as like adapting open weight LLM's. Um we do CPT, SFT like all of those stuff that a lot of the frontier labs are doing as well as like other sort of our competitors are doing and we try to kind of make sure that that we're building the best uh music recommender system possible for you guys.
So my my background is as a machine learning engineer. I used to work at Twitter and I live in London so if you're around for a coffee chat I'm very happy to after this talk or in general like talk about this stuff because it's really cool. Um so three things I'm going to talk about today. Uh the first thing is going to be about foundational user modeling. Uh so, what is what what is that? That's essentially us trying to understand you, the users.
Um these are like the three key aspects that we think are like weighty key aspects of the future of personalization in the era of LLMs. So, there's the the user modeling component. Uh then there's the content side. So, how can you kind of teach LLMs about the content that you have, the catalog that you have on on Spotify or whatever your sort of um your uh platform is. And then the last thing, once you have these two like pieces of the puzzle, are building these both of these together into like something which is steerable and personalized.
So, as as sort of steerable as possible. And I'm going to talk a bit more about that. So, generally like you can think about it as going from like sequences of actions to to vectors. Uh and then once you have the vectors, you can go to tokens. And then once you have the tokens, you can actually combine the vectors and the tokens with the LLM that you have to get like like a pipeline where you you have something which is as personalized as possible.
Um So, a bit about Spotify. Uh for those of you who have haven't used it. Uh we have like about 750 million users right now. Uh MAUs. Uh we have like a catalog of like 100 million plus tracks. Uh we have I think about 350 I think it's more like 400,000 audiobooks now. Uh millions of podcasts and like a lot of video episodes as well. So, we're we're kind of uh more and more creators are sort of switching to video as as a modality.
And we we definitely kind of making sure that we support that. Um and we're in about 184 markets. Um So, as you can see like we we have a lot of users, we have a lot of data and content. How can we combine all of that to build something which is which is as useful for our users as possible. Um the way we do that is we obviously we've been using machine learning models for at least like a decade if not more. Um you might have heard of like Discover Weekly or you're probably a user of that.
Um that's been around since I think 2015 uh back when I was in grad school and that was one at that time that was like for me like one of the coolest products uh with like my interactions across like the text sort of stack that I was using at the time just because like it's something which which is like unique to you, it keeps changing. Um and over the last decade or so we've kind of added more and more personalization in products uh to the to the app.
Um so we have like, you know, we have a bunch of verticals, we have a bunch of new product offices. Uh we have this thing called the AI DJ where you can kind of talk to it and it kind of recommends you stuff or it plays stuff for you. Uh we also have like a prompted playlist where you can actually prompt the model uh like Spotify's model and it kind of generates like a custom playlist for you based on your prompts and as of this week it also supports podcasts.
So if you want you can kind of just prompt it and it'll create like a like a playlist of episodes for you based on whatever you're looking for. Um so that is sort of the future that we're heading towards where users have steerability. Uh users can kind of talk to Spotify natural language. And uh now we also have something called the taste profile. So this is this is only supported like in a few markets but this is going to be expanded uh later this year.
And the idea is that we want to expose what we know about you. Um and then we want to kind of let you choose which part of that you want us to kind of keep, which part of that you want us to forget and you know, just allow you to kind of have as much control as possible. Um as we kind of work on this like just sort of some context for those of you who've not worked in the space of like recommender systems and machine learning.
So uh what we call trad wrecks which used to be like the sort of predominant uh sort of paradigm of building recommender systems up until a few years ago uh is essentially like this multi-step pipeline where you have like a massive catalog of items. Um, you have this candidate generation step which kind of reduces that item space to from like millions to like a few hundred. And then you have like a ranking stage and sometimes like you have multiple rankers which essentially bring that further down and then give you like the final list of whatever like you know, your top songs that we want to recommend to you.
Uh And we use this like across different products. So we have like home shelf ranking and we have personalized playlists and search and podcasts and like ads and like a lot of other stuff. And this is generally every team like every product has its own team that has its own model. So it's kind of like spread across different groups of people and some models are better than others, some have different features. So we're kind of moving away from that sort of siloed model of working towards this single unified model uh which is which supports uh similar to how LLMs work, which supports like an LLM backbone and which allows you to kind of uh allows you the user to kind of steer it towards the sort of recommendations that you want.
Um, and one of the key components of that so uh is is the user modeling part. So that is like the team that I work that I work with and what we do is we build user embeddings uh which are essentially representations of vectors that that tell Spotify about the user's taste across like all of the sort of history that we have on you on you the user like across the different sort of uh sessions that you've had with us over the years.
Um, and that becomes like the foundation of all of the models that are downstream uh that that actually um recommend stuff or allow you to search for stuff. Um, and these models are like very complex so they they the embedding model is we generate embeddings for like a billion plus users because we have we have a lot of users in general uh that are also like MAUs or that are not MAUs. Um, so we do that every day, so it's it's like a massive uh pipeline.
It's very expensive. Um, and over the years, like we've kind of moved away from having these sort of generalized user representations, which was like the predominant paradigm in machine learning, where you had these uh these models like this in this case like this was a paper from our team last year uh where we we kind of publicly spoke about the user embedding model that we have, which is generally um like at that time it was like an autoencoder model.
If you're familiar with that, what that does is it kind of takes all of your features, it compresses it down to like a small vector, and then recreates your features from that. And that sort of compression decompression process allows the model to sort of learn about you the user and just um represent represent user interactions in the form of a vector. So this is fairly standard stuff like that that a a lot of like folks do in the NLP computer vision space as well.
That is also kind of aligned with the way things were in recommendations. Uh we're kind of moving away from that towards uh the foundation modeling side. So now we kind of have this single sequential model, which as as you can see like with the the so the whole industry moving towards transformers and the whole shift that's happening across uh not just like the regular tech industry, but also the the companies that uh they use recommender systems and build recommender systems as as sort of their main uh bread and butter.
Um, so we're also a part of that sort of one of those companies, and we we we're also moving towards using transformers for for the for that kind of stuff. And the idea is that you want to have uh the the users interactions as a part of the prompt, so it's it's kind of like the context like I guess when we talk about context engineering, this is the context that we're talking about. Um, and then there's obviously the request con- context, there's like the query, there's the product surface, etc.
Like all of that stuff. And then there is the item that that you're recommending. So when you add all of this stuff, and then you put like a transformer layer and like a bunch of heads and all that stuff, and you you train it over with millions or hundreds of millions of users data, what you get is something really cool. Um what you get is something like this. Um So, this is an image which is from one of our like newer models which what it what it essentially shows you this is like a compressed version of what the model is learning.
Um in this case we have tracks in the blue and we have like episodes which are like podcast episodes in pink and then users so that that the one in green is me and some of the other green ones are like other folks in my team. Um so, that kind of shows you that we're kind of doing this cross content modeling where embedding users, tracks, and episodes in the same sort of content space or or I guess embedding space. Um and you're able to kind of visualize like I guess from from this image of like how on the hypersphere where you live alongside like different pieces of content and what is close to you, what is not close to you, and how can you kind of explore that that neighborhood region and how you kind of where you live you know, contextualized by where your friends live and things like that.
So, this is sort of like a visualization of like what the model is learning and as you can see for me like I am a machine learning engineer I care a lot about keeping up with with what Entropic is doing and what's happening in the tech industry and all that stuff. So, my specific embedding is really close to this big big tech podcast and on the right you can see that like that the the point where you see like those lines spreading that is me and then the ones in pink are like streams for tracks and the the blue ones are streams for episodes.
Uh and or sorry, it's it's the reverse. Um and you can kind of visualize like you can contextualize whatever you're listening, whatever you're like not listening to, and how how how does the embedding space look like for users. So, these models are like really smart um and the moment you give them uh information about the user uh they can they can kind of learn to put everything together in in in in a the space which is really cool.
Um this next part is more about catalog understanding. So, now that we have the users, we understand the users, we have a model for them, how do we understand the catalog? Right? Uh so, catalog understanding generally uh there's like a number of ways that you can do you can understand the catalog. Like the most common way is you train like similar to the to the user side, like you you train a vector to understand the content.
So, you have a vector that represents the item, whether that's a song, or it's an artist, or it's a podcast, or episode. Um and you have that vector. Um and then alongside that you have the user vector. So, that is like the Spotify knowledge. So, that's what we know about the the content or the sort of the different entities that we're dealing with. Um and then you have uh the world knowledge. So, that's coming not from Spotify, but it's coming from these open-weight LLMs that we're working with.
Um so, models like Llama or Gwen or like other sort of open-source models. Um what we do is we fine-tune those models, and then we kind of embed Spotify's knowledge into those models through uh something that I'm going to talk about. And uh what that does is it gives you steerability, it gives you better recommendations, it gives you explainability. So, there's a lot of stuff that you get for free when you use language models for recommendations.
Um now that there are like tradeoffs here, so that the model does end up like forgetting stuff, like catastrophic forgetting is an issue. Um but generally from what we've seen, like these models are really good at combining world knowledge with whatever sort of the knowledge that you have from your platform, and building something which you can kind of holistically use for for recommendations. Um now, the stuff that I was talking about where on the previous slide, um how do we actually teach these LLMs about the content?
Um let's start with that. So, the way that we do that uh is using something called semantic IDs. Um semantic IDs is like a fairly new concept. It it I think there's there was a paper from uh Google uh a few years ago which sort of introduced this concept in the context of YouTube. Um and what it does is like when you have a vector that represents a piece of content, so that's like a track or an episode in our case, what we do is we tokenize it similar to how LLMs like kind of we we tokenize words.
Um and what that does is it kind of compresses that massive like let's say a thousand dimensional vector into like four or six tokens. And that allows us to uh really use those tokens to train the the LLM in the way that like LLMs are usually trained. And it allows the LLM to kind of auto-regressively generate the next token or in this case the next token is not a word, but it's the next song or it's the next episode that you're going to be listening to.
Um so that's that's kind of what we're doing. Uh we're post-training or we're I guess we're like uh continually training these LLMs with Spotify's sort of data that we have about the catalog. Uh we use semantic IDs to compress the like the vectors into semantic or sorry, the vectors into semantic IDs. And at the bottom you can see that we have examples of like Ariana Grande and Bruno Mars. So uh we represent them as six tokens.
So those numbers are actually like token IDs. Um and the first two tokens for both of them are shared because they're both like pop artists and they're both like they both share something uh between them, but then the other tokens are different because those tokens sort of represent like more niches. So it's kind of like a hierarchical uh structure where you're you're kind of compressing the embedding into these six tokens and there's there's a hierarchy to it.
And that allows the model auto-regressively generate like the next artist or the next song that you're going to be listening to. Um so this kind of just overall like this kind of emphasizes how we do this. Uh we kind of we use the user context. In this case we have like a user who's Italian. Uh their listening history which is tokenized. Uh we send that listening history. Uh we use that in the training data. So we we kind of teach the LLM how to talk with semantic IDs.
And uh that is like the domain adaptation that I was referring to earlier in the slides. And then the final output is essentially generating the next item. So, whether that's going to be an episode or it's going to be a track, like what is this person like listening to? Um This is sort of like one example, like taking the example of the Italian person who maybe listens to an episode, like an Italian podcast. This is an example of a prompt that we give that model.
And you can see like the prompt on the left has the Spotify URI, which is like our representation of the of of the uh the item, in this case the episode. We convert that into a semantic ID, which is like the actual tokens that the the model is um attending to. And that is used to finally predict what is the next episode that the model that the user is going to be listening to. Um So, this is the catalog understanding part of it.
So, now that we have the user modeling part, we have the catalog understanding part, the next step is to essentially assemble all all of these components to form like a single steerable personalized generative recommender system. So, this is sort of moving away from the traditional rec- recommender system model to to this generative model. Um This is the sorry, the product that I was talking about that we launched just a few weeks ago.
This is called the taste profile. The idea is that you have some piece of text that represents what the the user is. We we expose it to you the user, and then you're allowed to kind of tell Spotify by chatting or by kind of sending messages or I guess adding some test text. So, maybe you want to start listening to like Justin Bieber more, or you you don't like this specific podcast that's being recommended to you. And what this does is allows the model to like that that data that edit or is going to come back into the generative model, and it's going to upgrade like the model's kind of ability to to sort of understand you and recommend stuff that's that's better for you.
Um so essentially like we have the content piece of the puzzle, but we don't have the user piece yet. Uh and the user piece doesn't come because ultimately these models are trained on like a limited amount of training data. You cannot train them on every like 750 million plus users that we have. Um so there is going to be some sort of level of like collaborative filtering. So the the the model is going to kind of generalize hopefully, but it also needs to be personalized.
Um the way that we do that is again coming back to the idea of user models. Um so uh you have the LLM. Uh you have the user representation. What you do is you project the user representation into the space of the LLM. Uh and what that does is it kind of allows you to create what's called a soft token within the model and it's essentially a token that represents the user, which is obviously contextually changed depending on on the user that that we're kind of generating this response Um and that kind of allows the model to be personalized.
So that is kind of like the final piece of the puzzle where you have this vector projection, which is like you have the regular LLM and then you have a user vector, which is projected to the space of the LLM. And that gets kind of inserted into the prompt and then finally when the model is actually like generating a recommendation, um the model is personalized because it kind of has the context on you like whoever uh we're kind of generating the the recommendation for.
Um these are kind of like the early results that we've had like we've seen some pretty positive results on our internal metrics. Um if you use Spotify like if you use the next episode sort of if you listen to like podcast on Spotify like this is something which is actually productionized now. So you if you're getting a recommendation, it's it's coming from from something like this. Um and that's that's kind of like the final piece.
So uh we have the embeddings which represent the users. We have the semantic IDs which represent uh like a compressed version of the content. And we have this soft tokenization approach, which allows you to project like users into the token space of the model. And this kind of moves away from the traditional recommender system model, and it kind of moves towards this sort of sequential modeling uh framework, uh which is something that we're kind of very excited about.
And uh that's Yeah, I think that's that's definitely uh going to be like really exciting as we go forward and we build this more more and more into all of the recommender systems that we have. Uh and we're very excited to kind of put this out there uh pretty soon. And yeah, uh that's that's my time. I'm I'm really happy to connect or like if you have any questions, feel free to reach out after the talk.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.