Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Discover AI · @code4AI
Words
4,636
Runtime
31:00
Speaking pace
150wpm
Reading time
19min
150 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Hello community. Welcome back. It's so great that you are here. Now you know that the whole development in artificial intelligence is just amazing and you're not going to believe today's video. Let's talk about harness as a language. You know, we have a new idea. We have the idea that we do not need a complex harness for our AI agents at all. And I know I know what you're going to say. You
75 words, the words spoken in the first 30 seconds at 150 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 316 |
| Average words per sentence | 14.7 |
| Longest sentence | 74 words |
| Questions asked | 34 |
| Sentences containing a number | 29 |
Most used terms
Filler phrases
35 in total: like 13 · I mean 7 · you know 6 · kind of 3 · literally 3 · right? 2 · uh 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
Hello community. Welcome back. It's so great that you are here. Now you know that the whole development in artificial intelligence is just amazing and you're not going to believe today's video. Let's talk about harness as a language. You know, we have a new idea. We have the idea that we do not need a complex harness for our AI agents at all. And I know I know what you're going to say. You say this is nonsense. For the last month we developed complex harness structure just to have a better AI performance.
Well, you're not going to believe this, but it turns out that the authentic behaviors of our AI system like a memory structure, a dynamic memory structure and a self-improvement, a recursive self-improvement or a regularized self-improvement. And you might say, "Yeah, harness. No, this is the intellectual cathedral of our year 2026. This is the masterpiece." It turns out that these things like memory and self-improvement do not require external systems like harness complexities in AI systems.
Because it turns out that they naturally emerge from a raw computational expressivity. And if you are now without any words, I do understand it. Believe me, I do understand it because this is just this is just crazy. Now we have today a piece where some authors claim that if you give an LLM access to a Turing-complete environment where everything, and I mean its prompt, its history and its tool is just a computational variable that it can manipulate with code and allow it to call recursive version of itself.
Let's call this recursive version simply subagents. This LLM can build these capabilities like memory and self-improvement on the fly without any complex harness at all. Now, you might remember that just 1 day ago I showed you in the new regularized RSI of the medical RSI from Google and Stanford, at the end I had this simple idea, "Hey, is this really true? If we build the right evolutionary environment, the AI will build the workflow itself?" And I mean, just 48 hours later we have a new paper and this show us that this idea is indeed kind of the solution for the next generation of AIs.
And you're going to ask me, "Who who published this?" Now, it turns out it is MIT, Massachusetts Institute of Technology, September 22nd, Harness as a language. A minimalistic agent framework with a maximal expressivity. Okay, let's look at this. Now, how to introduce this? You remember in physics in mathematical physics, the second semester if you study physics, theoretical physics at your university, you have pure mathematics and you have the principle of least action in physics.
And you understand, you don't need a million rules to explain how every different planet in our system or any exoplanet here in our galaxy, how they orbit, how they behave, or how light bends in a gravitational field. You do not have provided data set of hundreds and thousands and tens of thousands of example for the human brain or an AI to understand how we can calculate here the the orbit of a planet. You just need one fundamental rule.
You need one fundamental abstraction. You need one fundamental mathematical formula and the complexity and the beautiful behavior of a highly dynamic system emerges, like the dynamic of our solar system. We don't need training examples. We just need the abstracted mathematical formula. It turns out that this this applied to computer science might be the next breakthrough because given an LLM enough raw expressive power, the ability to see everything as a variable and the ability to spawn sub-regions, it will the LLM will invent memory.
And it will invent a self-improvement out of the LLM in a loop itself. No external harness, no vector databases, no graph databases, no complex structures at all. But the capability was never in the harness, it it turns out. Maybe it was in the computational substrate all along. And if you're going crazy at this moment, believe me, I do exactly know what you feel right now. Now, this the idea of this new scientific preprint here by MIT on Jazz is the following.
Hey, let's look for the fundamental physics of AI computation. So, they want to strip down the AI complexity of an AI agent down to a single primitive and this is here the command invoke. Now, MIT has two very simple pieces working together in a loop. And please, this is just a prototype, but it shows the power. So, you have the LLM at the core. This is classical in the agents, LLM at the core. You are a fable, whatever, your Opus 5.5.
And then you do not have a complex harness. You just have a totally naked Python execution environment, a very specific one, equipped with one special function. And this function is invoke. That's it. This is it. This is the new AI system. No database, no nothing. This is the entire harness complexity. It is incredible lightweight. It is incredible fast. Now, this new AI framework, let's call it Jess, like the MIT does, does not do any thinking.
It doesn't organize files. It doesn't search databases. It doesn't grade AI work. It is not nothing at all. So, what the hell does it do? This Auto Lisp harness has a simple job, enforce the loop. And the loop looks exactly like this, and the MIT defines this. It has a read, an evaluation, a print, and the loop itself. And you might say, "That's it? That's not possible." Now, let's think back, because if you studied computer science, you know that in 1964, yeah, in a different millennium, the expression read-eval-print cycle is used by Peter Deutsch and Edmund Berkeley.
What a surprise. For a 1964 implementation of Lisp on a computer recalled at this time a PDP-1. And you might say, "You're not really going to talk about computer science in 1964." Oh, yeah, we do. And here from Wikipedia, you know, for historic data, I love it. And they say, "Yeah, this our EPL, the user enter one or more expression, and the our EPL evaluates them and displays the result. The name read-eval-print loop comes from the name of the Lisp primitive function, which implemented this functionality.
The read function accepts an expression from the user and passes it it to a data structure in memory. For instance, the user may enter the S-expression here plus one two three, which is parsed into a linked list containing four data elements. The evaluation function takes this internal data structure and evaluates it. In Lisp, evaluating an S-expression begin with the name of the function means calling that function on the argument that make up the rest of the expression.
So, the function plus is called on the arguments 1 2 3, yielding the result equal 6. The print function takes the result yielded by evaluation and prints it out to the user. If it is a complex expression, it may be pretty printed to make it easier to understand. The development or the development environment then returns the read state, creates a loop, and terminates when the program is closed. Why we did this? It facilitated the exploratory programming and the debugging because the programmer can inspect the printed result before deciding which expression or what expression to provide for the next read.
Now, this new framework builds on exactly this. So, they build now a Python agent loop where invoke allows the LLM to write code into a read-eval-print loop. Now, crucially, the LLM's entire conversation history and the user prompts are passed as a standard Python variable. And you might say, "Yeah, and what is special?" Yeah, and they call this just Now, let's look a little bit closer to mathematical foundation because now it gets really interesting.
We are here in the year 1930, and you think I'm joking. No, we are in 1930, and we invent or we extend now the lambda calculus. And the lambda calculus, you read in the history of mathematics of human mankind, the lambda calculus is the purest mathematical representation of computation. And MIT took the lambda calculus with a new primitive which they called invoke. And unlike a normal function in computer code, the body of invoke is generated now dynamically and hold on to your socks by an LLM at runtime based on the inputs you provide for the query.
Now it turns out that this is Turing complete despite having here a real sparse syntax or lambda calculus built entirely from function. Remember invoke is a very specific a very special function in a very specific mathematical environment. It turns out it is Turing complete. It can simulate any Turing machine and perform arithmetic, logic, and recursion. And you know what we have mainly in AI calculation? Exactly. Now remember it's concept directly inspired from 1932 and 1964.
It inspired modern programming features like anonymous function and closer found in Python, Java, and yeah, you got it. Yeah. So MIT took this, built this, and said let's let's test it, no? Let's test it on a long horizon memory structure, no? And simply by instructing here the new agent to pass its history variable to a sub agent when its context windows gets let's say about 70% full process the call by the way a tail recursive delegation the prompt only jazz framework with invoke outperformed here one of the very latest specialized memory framework here letter.
Remember this was memory GPT but 8% in the performance and it costs now half as much as the memory GPT approach. Now even if you say okay, it just outperformed by 8% this is not more than 10% but it costs half as much. Now if you're not familiar with the good old times here from February 2024 and mem GPT this is now here. We introduce MemGPT a system that intelligently manage different storage tiers in order to effectively provide extended context within the LLM limited context window, right?
Later, it was 2025 and 2026, it was rebranded and evolved into the open-source agent framework Leta. Why do I tell you this? Here we have the experimental data, the result of MIT's experiment with Jazz Invoke. And here we have Leta, Memory GPT. Now, here on the left-hand side, you see here the first benchmark here for the long horizon. On the x-axis, we have simply the cost, the money, US dollars. Now, on the y-axis, we have the pass rate in percentage, right?
Now, look at this. It turns out that this new technology, Jazz Invoke, not only outperforms Leta Memory GPT structures, but it is also, I don't know, more than 50% cheaper. I mean, look at this. There will be a second test. This has a particular interpretation. I will give you this interpretation in a minute, but there's a second test where we have some other system, Ace. And it shows that even with a Mita Honesty structure, Jazz Invoke outperforms Ace in a different set of benchmark where different capability of an AI agent, and it is much cheaper.
Yeah, if you're not familiar with Ace, here we are, March 2026 version 3. This is by Stanford University and UC Berkeley. Agentic context is engineering, Ace, evolving context for self-improving language model. And you say, "What a coincidence, I know." So, now having said this, having you given an idea and having you show the result that it really works, that MIT found something where it claims, "Hey, we have two different benchmarks where we can have the interpretation that it invents memory and that it invents here a self-evolutionary path without any heavy harness structure, just a pure Python environment with invoke as a very specific function, this is our new AI system.
And I sat down and I really had to start from the beginning. So, let's understand what is happening here. Let's say you summarize a brilliant physics lecture, now, from Richard Feynman, now, and say, "Hey, listen, we had a beautiful lecture more than 2 hours and you know, I summarize it, hey, we talked about the wave function of electrons, now." I mean, you lose all the beautiful mathematics. This would be a crime. So, this new Chess preprint, this new framework on AI solved this by using a brilliant concept from computer science that in the good old times was called tail recursive delegation.
So, what is it? I give you an example. Imagine a runner in a relay race, now. One runner has a giant backpack filled with the perfectly preserved lab notebooks with a video recording of Feynman's lab or of Feynman's presentation, eh? The runner is getting exhausted. They're reaching here the end of the track. Guess what? This is the context window of the AI machine. So, what do they do if they want to hand it over to another agent?
Do they sit down, quickly write a one-page summary, AI-generated summary, and hand the summary to the next runner? Turns out, no, you lose extreme detailed mathematical information. In the Chess framework, because the notebooks are just digital variables in a particular environment, the exhausted runner simply hands the keys to the original backpack to the next runner. So, they simply write a code, return, invoke previous history equals history.
So, they pass it by reference. So, the next one, uh our next agent or sub-agent pops into existence as initiated by the first AI system, completely fresh, beautiful-looking, great with a brand new track ahead of them, but holding the exact mathematical perfect unsummarized notebooks of everyone who came before them. So, in the paper, one agent delegates the 69 times, and on the seventh time, it perfectly recalled a fictional protocol from 500 tasks ago because the notebook was completely intact and had the complete full mathematical description.
Nothing was lost. No information was compressed, no information was summarized. We had access to the full thing. And we had no database. Now, to the question how does AI teach itself in this very strange new AI framework, chess? A top-level chess agent might think just by accident about as professional like a physics professor, it dispatches here the sub-agent to solve problems in environment called, yeah, let's go with AppWorld, now.
Now, if a sub-agent failed, the professor doesn't just look at the final grade. The professor takes here the student exact notebook, the complete history of this agent or sub-agent, reads the step-by-step scratchpad where the student messed up the code or whatever, and says, "Aha, now you understand what you do." Look, you used here this tool in the incorrect way. Then the professor rewrites the tools instruction and gives the new instruction to the next student to improve the learning.
You see, this is a simple but beautiful loop. But have you noticed there are no rigid hardcoded learning frameworks at all. It is just one mathematical function in woke calling itself inspecting its own sort process and evolving its own code. If you want to go into biology, this is like a genetic code. This is like a living organism rewriting its own DNA on the fly. And I know what you're thinking. And I mean, I was exactly there and I said, "Wait a minute.
If we were literally pasting in a massive thousands and thousands of pages long text history into an LLM context window on every single turn. I mean, this would be a disaster. This is exactly what we faced 3 years ago in the eye. And here's now the absolute genius of Jess. Guess what MIT thought about this? They don't put the history in the context window. They put it in the computer's RAM. So, the Jess agent lives inside a Python code environment and REPL First agent is running out of breath, the second or the context window is getting full, it simply passes on its history to the next sub agent.
But it passes it on by a reference as a Python variable, not as a complete text string. So, in the new agent prompt, it doesn't see a million words of text. And the serializer simply shows a tiny placeholder that essentially says, "Hey, previous history, a Python list containing I don't know 5,000 history entries." And by giving the AI now total programmatic access to its history as a variable itself, this new Jess framework doesn't force or doesn't need a specific heavy architecture.
It just provides raw tools and lets the LLM write a simple computational efficient for loop to get exactly what it needs exactly when it needs it. So, if If want, we don't need any heavy memory function in the harness. It is incredibly efficient because it lets the neural network do what neural networks are good at, reasoning and writing code. Plot code. And let's the RAM, the CPU do what they are good at, storing and searching massive strings.
So, have we reached the state of perfect harmony? Now, remember, the massive string isn't in the AI's short-term memory in the context window. This massive string is sitting quietly in the background memory of the classic computer in the RAM. So, the AI is just holding, if you want, carrying you the library card to a particular piece of information, but not the whole library content itself. The agent doesn't try to use its neural network to read the history.
Instead, the LLM literally just writes a five-line Python script inside this environment, and this is one of the examples that you see now on your screen. So, this means the LLM delegates the brute-force search to the standard Python interpreter. And guess what? Classical CPUs are exceptionally good at finding substrings in massive blocks of memory. They can search megabytes of text in microseconds, and it costs virtually zero energy and zero API tokens.
Here on the left-hand side, you see now a screenshot from example here from this MIT publication explaining in vogue. So, you have to use a prompt, available tools, and and the action history are not treated as any special magic elements now in the classical heavy harness of an agent. Here in this new interpretation, in this new, if you want, principle of least action applied to computer science, those elements are just standard Python variables passed into a particular REPL loop that the LLM can manipulate.
This is it. Nothing else. So, in just invoke the function invoke acts as a function whose implementation is provided each time it is called. And who does it? Well, it's provided by the LLM in a real-time Python loop. So, everything, all named inputs to invoke and the REPL history itself, are variables that are now available in the REPL. You have here second screenshot from the MIT where they show you the consequences of this particular invoke definition.
And look at this. You see here that we have here the definition of recursive sub-agents. And they are by default here. And I will tell you some details about recursive sub-agents or default at the very end of this video. There's a kind of a surprise waiting for you. Yeah. A closed agent loop arises simply from tail recursive invoke functions. Now, they provide you with a lot of details, a lot of prompt structure, a lot of guidance.
Look here. You have everything that you need to build it yourself. You have the complete GitHub for the evaluation, for build-up. MIT really provides you the complete code. Here a screenshot here from the continual self-improvement where they define the sub-agent context, the sub-agent prompt, the sub-agent tools, observability, metric, traces. Everything is there for you. So, you might say, "Heaven's sake, we really have now invoke?
This is a very special function written at runtime by the LLM, and we do not need a harness functionality anymore?" Yeah, exactly. This is it. Let's look at the results. This is now interesting. This is here the first benchmark I saw over the shoulder you now here the numerical data. So, we just go here with letter the letter agent because this is one of the most recent and most powerful competitor, and you see here if you look at this here my far recall here the score, we go from 67% now with this new just invoke methodology to 73.6%.
So, if you want this is here the empirical proof by the MIT of their memory claim, no? Because it compares it a minimal just invoke setup against letter the state-of-the-art memory GPT agent designed specifically for long-term memory. And here the result prove that just handing here the I it's previous history variable scoring 73.6 mathematically and financially beats here a heavily engineered external database system letter.
So, this is quite impressive. But then we have another test. This is here app fold, and this is here according to the MIT the empirical proof of the meta learning claim of this new framework just. And here we go with ace here this one of the last lines and here just invoke the very last line, and we just compare here the performance. So, we go from 69.9% to 74.2%. And yeah, the cost are significantly cheaper. And you don't have to build this heavyweight harness structure.
And you don't have to pay for this extreme amount of tokens. So, there's something that I really like in this paper by MIT. Outperforms ace, you got it. Now examples. If you still say, "Hey, what is invoke?" There's a mathematical view and there's a computer code view. So, let's Here you have on the left-hand side on your screen, you see here the example given by MIT itself here. What is skill definition here? You have here the invoke command here with task instruction skill eval.
Beautiful. But if you are new to this, well, let me give you a simple example. Think of a standard computer function like a recipe, no? You have the function, hey, make breakfast, and your ingredients are eggs and bacon. The code inside this function tells the computer exactly what to do. Turn the stove on, fry the bacon for 5 minutes, scramble the eggs, and done, eh? The recipe is set in stone. Invoke is now a magic recipe box.
When the computer reaches you this particular function, invoke, with particular ingredients and particular goals, there's no code written inside it yet. It is completely empty. It needs an LLM. It needs an intelligence that for particular specific task, my human query, for example, code needs to be written into the function, invoke. So, you see, at the exact moment the computer runs invoke, an invisible chef now in this kitchen, and this is our AI LLM, looks at what you dropped into the context window, let's say the goal is to make French omelet, and the ingredients are eggs, butter, and herbs, and the LLM then, with its parametric knowledge, writes now code.
An LLM can write code, think about Claude Code, and writes the recipe on the fly inside this box of this very specific function, invoke. And it then executes the newly invented recipe and hands you the omelet. Done. Do you see the parallel to one 3 days ago? I had to do this video on Jeff, fast decision, but confident mistakes, where we talked about an abstraction level, where we said Jeff is kind of a very reduced, heavily reduced, AI model based on a very specific transformer architecture, but it has only a simple task, and it needs in its input certain option where it will output you a probability distribution on this certain option given here my particular human query.
So, this was if you want a very specific expert system that was operating here on a subset of an intelligence only do if you want decision models where you provide a probability distribution on options. This new methodology by MIT is on a completely different level of abstraction. Because it can do almost everything. Now, you remember the beginning of the video I showed you that there was this statement by MIT recursive sub agents are the default.
Let's talk about this. And let's talk about why this is amazing and it could be something amazing for AI in companies, in offices. A simple example. Imagine you have a very special button on your desk. And no, it's not for the Coke, but it is for a temporary agency. So, this is exactly our mathematical function invoke that you understand here the analog in this example, eh? So, when you press this button on your desk, a brand new temporary worker drops from the ceiling into the chair next to you.
And because you're using the mathematical function invoke, you don't just give the temporary worker any verbal instruction. You just literally hand them the things from your desk to put on their desk. And you say, "Listen, here's the folder of files." You pass over the data variables. And you say, "Hey, here's my calculator." You pass a tool. And you say, "Hey, here is a documentation with what I need to do." So, you pass over the task.
And yeah, you already see where I'm going with this. Now, the temporary worker goes to work. But wait, the temporary worker looks at the folder and realizes, "Wow, this is too much work for them. This is a too high complexity. It's not able to solve. But because we are now in the same fractal office structure in the same mathematical space, the temporary worker has also a desk and on its desk is also a temporary agency button on their desk.
And when if they press this button and they invoke the mathematical invoke function, they will spawn now a third worker and hand that third worker half of the files or and the same calculator and you get the idea. This is it. So, we do have an AI system which I call the perfect office, but just think about it, eh? If the workload is extremely high, the I will generate let's say it very very un-mathematically, it will generate copies of itself.
It will generate additional workers. They do not need a harness structure. They do not need databases or whatever. But they can generate copies of themselves on demand. And then at the end of the day, they are all deleted. The work is done and great. You go back to a single agent. I mean, the idea by MIT to do this without any complex harness structure is simply It shows you if you have to combined knowledge of computer science from the 1930s, from the 1960s and you apply this to something where everybody is running to design the most complex harness architecture and MIT says, "Let's think about it.
Let's make it a smarter AI model. And let's see how far we can abstract everything away. And let's see what is at the real core of the mathematical abstraction of an AI computation in computer code. And it turns out it is a simple function called invoke. I hope you had fun with this video that was a new information. Read the paper. Of course, I did not went into the complete deep mathematical interpretation. I will leave this to a beautiful afternoon to you where you will enjoy this paper with a beautiful cup of tea.
Looking forward to see you in the next video.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.