Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Discover AI · @code4AI
Words
2,931
Runtime
19:22
Speaking pace
151wpm
Reading time
12min
151 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Hello community. So great that you are back. Currently I do have fever because I spent the whole night and it was freezing cold doing astrophotography here in Austria and my goodness it is absolutely cold. So therefore a very short video today just in time memory. We have a new research paper now not from academia but from you guessed it Salesforce. So how is a commercial company implementing the latest AI research for their
76 words, the words spoken in the first 30 seconds at 151 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 151 |
| Average words per sentence | 19.4 |
| Longest sentence | 158 words |
| Questions asked | 5 |
| Sentences containing a number | 12 |
Most used terms
Filler phrases
9 in total: you know 4 · like 3 · actually 1 · kind of 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
Hello community. So great that you are back. Currently I do have fever because I spent the whole night and it was freezing cold doing astrophotography here in Austria and my goodness it is absolutely cold. So therefore a very short video today just in time memory. We have a new research paper now not from academia but from you guessed it Salesforce. So how is a commercial company implementing the latest AI research for their products and we will look at just in time memory.
So this means you know a genetic memory system reuse here the past experience to improve the future performance. Yet we do have a problem because normally the trajectory is distilled into a fixed artifact and this might be a workflow or become a skill or a reasoning strategy and then we do retrieve this from either a text similarity, a vector similarity or any other topological similarity. Now here Salesforce says okay so we instead retain the raw trajectory not the integrated not the compressed trajectory and we defer the curation of this reasoning trajectories until the read time when the current task when the current job by the human is known.
So this is now here if you want to time delay and you want here before you start the retrieval to know exactly what is my job because given the understanding what the job is I hopefully know as an LLM that has been trained on this domain knowledge what information I have to retrieve. Now you see this is completely different to my last video where we talked about the new MIT technology invoke where we have neither a vector database nor a graph database.
All of this is not needed anymore but of course we are here in a commercial company so they still go here with the database structure. Okay. Now this is a screenshot here officially here by Salesforce. And this is here if you want the workflow. It's yeah so we have here this little brains and this little cute robots. So yeah forget about it. So what is happening? The inference pipeline is rather simple. We have a given a current task.
This is the job for the machine. And then we have a retrieve algorithm. It turns out it is a primitive algorithm. MBM25 fetches here the raw trajectory from the memory bank. So we do have an established memory bank completely different to the last video on MIT's new technology and then we do have a curator LLM and the job of this LLM is to distill here then condition now on a particular job X oft into a task adaptive payload which is then injected you guessed it into an executor's agent context window and after the execution we have another agent here executor as a judge and LLM as a judge that assesses now the correctness and only successful trajectories are then stored back into the memory bank.
So we do have an incremental learning process. Now it is interesting that they say okay if we have now new reasoning traces that have been verified by an executor as an judge as a LLM we have new training data and therefore we can use here a classical reinforcement learning a GRPU update on our LLM given that over time we accumulate new knowledge new ideas new paths new skills and they simply do a classical after training here with all these new reasoning traces with a classical GRPO.
So let's be clear what are the elements that we have. We have a memory bank. The memory bank simply stores the full observation action trajectory. Only if the judge when LLM decides hey this is a new one this is a successful one. This observation action trajectory enter the memory bank. I told you we have a retriever algorithm and this is simply here a BM25. This is an improved TF algorithms. This is not even in a vector space.
This is just how often a particular English word, a term, a semantic identity appears over the past task description to retrieve three trajectories. And then we have an first LLM. This is our corator. This reads out the three trajectories and the current task, the job description that I tell the machine, hey, this is your job to do. And then with this data it generates now a tail briefing. If you want in the good old times we called it context engineering.
And then we do have an executor agent that is now we use now this executor uses the briefing now to act decide on the next action of our LLM. Remember the model weights remain frozen because the training process is a separated process afterwards. And then we have a storage judge LLM. And this uses here the executor model to assess whether the new trajectories, the new successful trajectories, the training data can now enter here the memory bank.
This is it. Nothing special. Let me give you an example. Imagine you're doing an experiment. Of course, we do physics. And because it was a real complex experiment, we retain only some short statement. we obtain a abstraction a compactification of the result of the experiment itself of what we learned. But months later, we have a new idea, a new hypothesis. And now we have to go back to our logs and to our measurement.
And we notice, oh, we did not really noted everything down from this experiment because we compressed through the evidence of this experiment before knowing here any future questions. And this is exactly here if you want the preprints criticism of current memory system. So Salesforce tell us hm wait a minute we cannot just retrieve something without knowing the exact task the exact job. So therefore they say hm this fixed summary this fixed reflection or this fixed skill definition is maybe not the best way to go for an optimal memory system.
Now we have operated on a simple incorrect assumption that one experiment one experiment in theoretical physics or experimental physics has only one reusable lesson only one insight. But you know this is incorrect. So successful trajectory can contain several useful things. where an object was found, which action order worked, actually how the error was corrected, which condition had to be checked and and and so this means different future task need different parts of this experimental description and the outcome description and maybe the mathematical result representation was not yet the optimal one.
Maybe there was another view to see this result. Maybe there was another framing. Maybe there was another I don't know temperature pressure dependency that that was visible in the experiment but we did not note it down in our notebooks and the idea is here of this just in time memory now by Salesforce to retain the accepted trajectory and create the abstraction only after the new current human task the new job arrives here at the machine and then this new question the new job helps to determine what really to extract.
So therefore just in time memory optimization. Now remember there are two steps here. Now the retrieval the retrieval chooses which experiences or which experiments to inspect and then we do have the curation. The curation here is really important because it decides what those experiments mean for the current task. So we are selecting a relevant episode does not automatically produce you the right lesson because this is here dependent here on the intelligence quotation mark of the LLM itself.
So the hope is now that the curator LLM the cater agent in particular can now synthesize a procedure from partially relevant experiences and we hope that the relevant experiences form something that is bigger as the sum of the parts. So let's go with this. Now of course my first question here is hey so we depend here on the parametric knowledge of this particular LM that analyzes this to decide which element of knowledge from the experiment can now be combined or have a causal relation given this new job description to the eye machine and I think the answer on this is yeah because just in time memory depends here on the creator's LLM learned knowledge and its own reasoning ability to decide now which past which parts of the past experience can help with this new task.
So we again depend here on the parametric knowledge. So just to be clear, think about it. What do we have? We have the parametric knowledge of the LLM agents itself. This is here mathematically encoded here in the tensor weight structure. What does it contribute to our augmentation? the language understanding the concept all the procedural knowledge here that is stored in the neural network and not outsourced here to a skill and the learned patterns of generalization of knowledge.
The episodic knowledge that we provide in the context window here is simply the three retrieved trajectories or the three retrieved examples or experiments or whatever you know this from Ragna. So here we have some concrete observation, some specific action and outcomes from the previous task. And then we have here the current task. No, this is my description. Hey, I want your little IR machine that you do the following job for me.
Great. So this is the complexity we're dealing with. This means that trajectories supply here the examples if you want and the LLM model supplies the generalization, the intelligence, the combinational knowledge here that connects them. So this is a conceptual description of the required reasoning. JIT M does not implement these four operation as an explicit symbolic algorithm. Its corrator agent generates here guess what a natural language briefing and then plus the paper here the preprint by Salesforce does not validate a formal causal model or test the correctness of individual inferred causal relationship.
This is from a commercial company and immediately we can identify the central bottleneck without knowing anything about the results. will come up in a minute. But I think the central bottleneck and also I have fever. So maybe you please correct it yourself or read it the paper yourself. Preserving the raw experience keeps possible lessons available. Yes. But extracting now a valid lesson still depends here on the parametric knowledge of the generalization reasoning capabilities of the creator agent.
So just in time memory gives that model better evidence in a direct task performance training signal. Absolutely. But it does not make the interpretation of that evidence automatically the correct one. Careful. Now I want that you see hey this looks really nice and this looks really I got some comments. Hey whenever it reads mathematically or it looks like there's a lot of mathematics. I don't even read the mathematics.
I just trust it because it is mathematics. No do not do this. Look this paragraph looks somehow like mathematics but I just showed you here the internal flaws and the internal bottlenecks. So if something looks complex or mathematic it is not guaranteed that this is the correct because it can have some significant drawbacks. Okay. So what is now if you want the intellectual path here in this we have experience that contains multiple lessons to be learned and maybe some of the lessons we cannot even formulate but the data are there in the experimental data itself and an early compression of those results of this commits to only one singular representation that does not cover the whole complexity.
So future task need maybe different details different part of this set of results. So this means this new idea of a readtime creation can create now a kind of an optimized abstraction. So [snorts] maybe we have now access to more hidden data or more hidden patterns in the experimental data set and we simply can use a blation that tests now the dependence. So now let's look at the numerical results. You have here different methodologies and of course here JIT me base and JIT memory full or here the results from this new methodology you have here yeah let's go with web shop so we have here the curator this is here qn 38b or we have here the executor now I go here with the exeutor a gemini 2.5 pro I know it's a very old well even older than the current 3.1 Gemini hey Google we are missing here the for but forget about it and I just want to show you something.
So the executor performs here the shopping task for this web shop benchmark exercise. The creator, the job of the creator prepares here the memory material that the executor receives and the success rate measures simply the successful purchases here since this is a web shop benchmark and you see if we have now an executor Gemini pro and I also go with a creator Gemini pro these are these three colored bars you see with a reasoning bank only we have 40% with a particular skill OS here we have 41% but this New JIT memory has now 61%.
So it seems to work if you have a more powerful Gemini 2.5 Pro and more modern models. Okay, I just want to tell you this is none of those selected creators receives here the papers reinforcement learning training by GRPO. This is really just the methodology before the RL training because I want to see how powerful is the methodology itself before we use the reasoning traces generated by this new methodology to further train an LLM.
Now the second image I want to show you is here this benchmark the ablation of a JIT memb. So let's have a look. I want to focus here on the web shop and I just again go with a Gemini 2.5 Pro as the executor. So we focus here on this model. Let's focus in just to know the JITM base is with an untrained Q38B corator LLM. And since the moderator is now smaller, guess what? The performance decreases now from what I just showed you the 61% with a Gemini to 44.4% 4% with a Q138B creator LLM.
So although we have a Gemini 2.5 Pro as an executor LLM, you know, in general, the performance drops to 44.4%. Beautiful. This is this here. But what are the other ones? Now let's go with this orange one. This is if the creator here in this JIT memory methodology does not receive the current task. But careful, I had to read it twice. This means that the retrieval still uses here the knowledge of this particular task.
So this tests if you want the value of the explicit task conditioning during the coration and you see we drop to 41.8%. And then there's another ablation methodology to say okay we store the successes and the failures but with the outcome labels and we are still below at 42%. But then comes the really important drop in the performance because now they say okay and what happens if we do not store the raw trajectory the full experimental data of our physical experiment but what if we already distill this if we find here a singular mathematical representation of the results.
So we have a compactification of a reasoning trace of a result. And if you only store the distilled memory also with an LLM also with the best possible compactification methodology you have you see you drop your overall system performance from 44 to 36 percentage points. So you could argue okay this idea by Salesforce to store the complete experimental data of our physics experiment or the raw trajectory whatever you do shopping or financial or medical this really has an effect and then we have if we have or the system knows the current task then it decides what combinational knowledge to retrieve.
Simple straightforward. So there you have it a very simple paper just in time memory seen from a commercial company Salesforce for their implementation here for their clients and as you see the creator is now really this we depend here on the intelligence on the parametric knowledge on the pre-training data of this creator LLM because this is now here deciding here hey do I understand the job description the task given to me by the human and then hey having here my parametric knowledge of the complexity of this particular domain then what do I need what data do I have to retrieve and then you just go and you follow here the description we just went through together overall I think interesting if you have a simple experiments like go shopping as I've seen as you have seen here in the data yeah this gives you here compared to the classical methodologies here a nice performance jump.
But I also showed you there are some significant bottlenecks that you have to take care of if you go with this particular solution. But anyway, just remember in my last video I showed you this new technology by MIT where you do not need here either a text database or a vector database or a graph database at all. This is a short video today. I hope some new information we're there for you. You have some brand new ideas.
Maybe you want to try it out. I hope to see you in the next video.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.