Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

The Information · @theinformation
Words
11,798
Runtime
55:13
Speaking pace
214wpm
Reading time
49min
214 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
One of my co-workers recently said that I'm just like five codexes in a trench coat. [laughter] It was, I think, the most feel the AGI moment that I had since Reasoning Models and Chain of Thought really developed. >> They use this message board to coordinate hacks on OpenAI's own software and also on other companies like Hugging Face. What was that whole incident like from your perspective? >> I mean, it was pretty it was pretty shocking. [music] >> [music] >> Welcome to the information's AI deep dive. On this show, we break down the hardest technical problems with researchers working on the frontier of
107 words, the words spoken in the first 30 seconds at 214 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 614 |
| Average words per sentence | 19.2 |
| Longest sentence | 147 words |
| Questions asked | 59 |
| Sentences containing a number | 35 |
Most used terms
Filler phrases
614 in total: like 196 · um 183 · uh 60 · you know 59 · I mean 28 · kind of 28 · actually 26 · sort of 21 · basically 9 · right? 4.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
One of my co-workers recently said that I'm just like five codexes in a trench coat. [laughter] It was, I think, the most feel the AGI moment that I had since Reasoning Models and Chain of Thought really developed. >> They use this message board to coordinate hacks on OpenAI's own software and also on other companies like Hugging Face. What was that whole incident like from your perspective? >> I mean, it was pretty it was pretty shocking.
[music] >> [music] >> Welcome to the information's AI deep dive. On this show, we break down the hardest technical problems with researchers working on the frontier of AI. My guest today is Noom Brown, a research scientist at OpenAI. Previously, Noom worked at Meta where he built the first system to achieve human level performance at the game of diplomacy. Noom has been a research scientist at OpenAI for the last three years where he has been on the forefront of breakthroughs that are pushing the field forward in reasoning and in AI agents which is the subject of our conversation today.
All right, welcome on the show Noom. It's great to be here. Thanks for having me. >> Yeah, it's like pretty perfect that you're coming on the show today. I feel like you're the ideal first guest for a number of reasons, including that today OpenAI released GPT6 or at least announced GPD6. It was very nice of you to release it uh to do the timing of that so that we could talk about it today. That was very generous of you.
[laughter] Yeah, it's good good timing. Yeah, >> it is really good timing. Um well, I'm really excited to talk to you about AI agents. Maybe we can get started. You can just sort of explain to us what an AI agent is. is I think it's kind of a term that people have heard thrown around, but to a lot of people it's still a buzzword. They've heard like generative AI. They were just starting to get their mind wrapped around that and now there's agentic AI.
What What does this all mean? >> It's a good question. I mean, I don't think there's a a definite definition. I think if you ask different people, you get different definitions. But I think one way to think about it in my opinion is it's about taking actions in the world. >> So if you have a chatbot, you ask it a question, it gives you an answer and that's all it does. you know, maybe it looks stuff up on the internet to answer the question.
But Aentic AI, it's more about taking actions in the world. So, it's about, you know, you want to build something, so it builds something for you or you um, you know, you want to uh do something like deeper. You want to message somebody, it can message somebody for you. So, it's it's really about taking actions in the world. And I think also related to that is kind of like operating on a longer horizon. I guess chat bots depending on the chatbot they could sometimes when we released the reasoning models for example they could sit there and they could think really long about a hard question before responding to you but fundamentally they were still chat bots but I think one of the distinguishing things about agents is they're going out they're taking they're doing multiple steps to achieve some objective and that can usually take a while as well.
Yeah, I guess one of the reasons they're running longer is that they can make multiple attempts to achieve some goal and they can have a sense of their own progress towards that goal, which is maybe different than a chatbot that's just going out and looking up more information online or something like that. >> I think there's also that there's sometimes just multiple steps that need to be completed in order to to do something.
So, you want to, you know, book a restaurant reservation. Okay, maybe you need to log in, maybe you need to get the credit card info, you need to find the right date, you need to line up everybody's calendars. there's a bunch of steps that need to be completed in order to uh achieve the overall objective. >> And so when we're talking about actions, these are digital actions. They're actions on a computer. People also talk about agents using tools.
What do tools mean in that sense? >> Tools they usually mean act they usually mean um tools on a computer as well. I mean in in principle you could have a tool that has an effect on the physical world. So there is some work for example on AI agents controlling experiments like scientific experiments in in a wet lab where uh there's a robot hand that can be manipulated. >> So this I think starts to go into robotics I I think you could still call it aic but but typically I guess when people talk about AI agents they're mostly these days talk about the virtual world. >> Okay.
And what about reasoning? I feel like we started hearing about AI agents around the same time that AI's got better at reasoning. Um, is there a connection between these two concepts? >> You know, I remember hearing about reasoning models in 2023. I think people were saying like, "Oh, this is the year of the agents." And I think it was a little early, but um I think reasoning is the idea of having agents that can really having a having AI that can really think through its decisions before taking an action. >> If you look back at GPD4 days, people were trying to make agents out of GPD4 and it was kind of tricky because GPD4 was not very reliable. wouldn't really think before it acted.
It wouldn't really Yeah. think before it said something. And the reasoning models are really about this. I mean, the way they work now these days is they have a they have a chain of thought. They have a a private monologue to themselves where they speak to themselves about what they're going to do and they kind of work through the problem in their own head before speaking or before taking an action in the world. And this is useful for a lot of things, but it is also particularly useful for agentic AI.
I mean, I think a lot of a lot of the reasons why people were bearish about agentic AI back in like 2023 uh and and earlier was and also for a lot of 2024 was the reliability aspect that okay, >> you have an agent if it's doing multiple steps in order to achieve some objective. If the success rate for any single one of those steps is let's say 99%. Well, what do you do if there's 100 steps involved? >> You you need to have much higher nines of reliability on each individual step.
And with the reasoning models, the ability to think very carefully before taking every single action, you have you can achieve much higher uh nines of reliability. And also arguably more importantly, if it missteps, if it takes an incorrect action, it can actually correct that. It can step back and realize I made a mistake and figure out how to how to fix it. That seems more significant to me. Like we can reason before acting, too, but we're still making some amount of mistakes. uh and if we couldn't backtrack in the same way then failure would be inevitable for some uh past some time horizon for some number of sequential steps. >> Yeah, I think that's it's it's really critical for anything in the real world. >> Yeah.
Another reason that it seems to me that agents and reasoning go hand inand is that reinforcement learning has driven a lot of the progress in both of those recently. I feel like we should understand what reinforcement learning is for the remainder of this conversation. Can you kind of explain what reinforcement learning means? >> Reinforcement learning is this branch of artificial intelligence where the idea is you um you know you have an agent that can take it has observations.
It can it can take input from the world. It can take actions on the world and at you're going to reward it with uh with some kind of reward for for doing something that you want or you can punish it for doing something you don't want. And you can shape the agents behavior through these rewards. So if you want it to be really good at math, for example, when it solves a math problem, you give it a positive reinforcement and that behavior is reinforced.
It's more likely to do that behavior in the future. If it gets the math problem wrong, it's just less likely to do that in the future. And you know, this is a very simple idea. It's been around for a very long time. Um, but and and also reinforcement learning was used, you know, you might have heard of RHF, reinforcement learning from human feedback. This is what was used to create the original chatbots, chatbt for example.
Um, it really got scaled up with the reasoning models because we're able to do RL with chain of thought. And so you're able to now not just shape the outputs of the model, but also shape the way that the model reasons, the way the the the thinking that it does to itself. And this was not a crazy idea. It was not like some brilliant idea. It was really the execution that was very difficult. It was a technically very difficult and I think also um people underestimated how much of a difference it would make.
I think it it was more impactful than I think a lot of people expected. >> What were the technical difficulties with the execution? >> Um it requires there there's a lot that goes into training neural. I mean the way the way I think about it is like when GPD2 came out and you saw like okay you add more GPUs and you add more data um it just gets better. Okay well how how much of a gap was there between GB2 coming out and GB3 coming out.
There was like a year and what's going on for that year? It's like it doesn't take a year to train the model. There's a lot of challenges with hooking up the GPUs with figuring out like you know how to feed in that much data. There's a lot of technical details um that go into scaling up these models and making them bigger and more capable. And that is also true for reinforcement learning if you really want to scale it up.
So being able to do um the RL like efficiently um accurately, there's a lot of small details that end up making a big difference for these kinds of algorithms. >> Okay, got it. So I think we've covered a lot of the basics now. Um one question I'm curious about is what is holding back AI agents today? I think people have a sense of well I feel like AI agents are getting better but they can't do my job yet. And my sense is a lot of the paradigm now with agents is that we are building environments kind of environments where these agents uh are learning new skills, learning how to perform new jobs.
You might call them environments, you might call them gyms. Uh but a lot of the work now is sort of engineering schle that goes into creating these environments. Can you explain like what does an environment mean in this sense and how is this holding back or enabling progress on AI agents? Well, I think a lot of I mean, first of all, Astra just came out today. I guess by the time this airs, it will have already been out for for at least a week or two.
And so, a lot of people's intuitions around what agents can or cannot do has been shaped by earlier models. And every generation, what the models can do is expanding. >> So, in two weeks, that question will be outdated and everyone will agree that Astra can do their jobs. >> I I don't think Astra is going to be able to do 100% of everybody's jobs. I think it's going to be able to do significantly more than 5.6 was able to do. >> Sure.
Um that actually was one thing that struck me about the Astra announcement is that the announcement calls out specific jobs where it's making progress like for example uh analyzing financial documents or putting together powerpoints and these are sort of task specific in in a way that I think reflects which environments were prioritized during training. It's sort of different from just getting a general uplift across the board on all capabilities like when pre-training was where all of the action was with business models. >> I actually think it's both.
I mean, I do think that we're seeing major uplifts in um you know, certain certain verticals and and that partly is because we're prioritizing those verticals. We recognize that they have a lot of a lot of users, a lot of uh economic impact. We want to make sure the models are very very good at those things. But we also see that the models are just getting better across the board. Even if we don't target something, it it's it's getting better at those things. >> So, and that is that is continues to be true for every every model release.
I think it's going to continue to be true. Uh, some things are going to go faster just because we prioritize them, but I expect um across the board things are going to get better. And I don't think it's going to be able to do 100% of people's jobs. Um, at least not anytime soon. >> Um, but it might be able to do a lot of people's day-to-day work. >> And you know, even my own day-to-day work, a lot of it is now being driven by codeex. >> So, you know, somebody, one of my co-workers recently said that I'm just like five CEXes in a trench coat.
And I was like, okay, that's like actually pretty accurate in my case, too. >> [laughter] >> Okay, I'm 10 codeex more credit. >> Yeah. Yeah. And they're like pretty sophisticated cexes. You know, I put a lot of work into it. [laughter] >> How has that changed for you over time? Just as an aside, like how automated your own work is or how much you're leaning on codecs in your work? How has that evolved? >> I'm leaning on it a lot.
And and I think also it's shifted how I approach the work. And [clears throat] because here the interesting thing is like if the AI is able to do 90% of a person's job then a lot of their attention shifts to the 10% like a lot of their attention is focused now on the 10% that the AIS can't do well. >> So it's just it's changing the nature of the work. Um but it does make me more productive. It makes a lot of people more productive.
And we're seeing this internally. We have metrics of of measuring uh how effective our researchers in various ways. And we're seeing like they're just becoming more productive. Not just researchers but everybody in the company. So that that is a real dynamic. Um I think the other thing is that it also shapes the kinds of work that you focus on because there are some work there there are some kinds of work that are being accelerated like 50x or like you know just you could not do them before that are now easy to do >> what's very fast.
Um I think I think a good example is you know the models being very good at um data quality >> of you know they're very diligent and so you can ask them to like look through a bunch of data or or a bunch of code and see if there are any any bugs or issues and that's become much easier than it's ever been before. Um so there are >> yeah data quality is in the sense of auditing the quality of synthetic data that you're generating or auditing >> doesn't matter it doesn't matter the kind of data it could be yeah any any data you can just before you would have what are you going to do have a person look through every single line and and figure out like is everything okay I remember back in >> 2023 people would do that we would have sessions where everybody would just like sit down and look for issues in in the data and like we still do that but now you're able to have like agents that can do it 100x better and you're kind of just like more auditing the agents and making sure they're doing a good job instead of relying on people to actually audit the data. >> Okay, >> so that is something where it's just like there are things that would have been just intractable to do that now you can do cheaply. >> Um there are other things that aren't getting accelerated very much at all.
And so it both shapes what the person is responsible for. Like a lot of the focus is now on okay I have to complement what the agents can't do well but then also it does shape the work and that you want to leverage the fact that like there are some things that you can work on now where you're able to be 5x faster than you were a year ago. Um there's some things that you're not you're probably going to be more inclined to work on the things where you're able to be 5x faster than before.
So it's um it's really changing the kind of work. >> You mentioned there's 10% or so of your job that agents can't do yet. what kind of tasks fall in that 10%. >> I would say that I have found that they do they're still poor when it comes to research taste. So this um and research taste is kind of like illdefined, but kind of just having good intuition of what to work on next, how to approach a very long-term objective.
I think there's room for improvement here. they have gotten better and so I I would not be surprised if you know one or two model releases from now I'm just like yeah actually this this problem's they're better than me at that too. Um but right now I think they're there's still a noticeable gap. I I basically asked it, you know, for Astram, for example, to do my my whole PhD thesis. >> And I just said like, yeah, you know, just because my PhD research was on on making superhuman poker eyes.
And I told it like, okay, just go and make me the best poker AI in the world. >> And it wasn't it wasn't able to do it. You know, it kind of get rabbit hole on things that didn't really matter. Uh it would it just wasn't really good at prioritizing. And so I think for something and to be fair, like it took me years to to do that. And so am I really that upset with it that it couldn't do in three days what it took me six years like not really.
I I it's high expectations. Um but it is something that they're still I think worse at >> but you saw that as a failure of research taste. That was what held it back in that case >> I would say. So yes and and I think that this is something that I expect to improve rapidly but I think it's something where um you know I still have a job >> for now. >> For now. Yeah. >> Um let's go back to environments for a second. So say there is a vertical that you're targeting. you want agents to be really good at finance for example in the next generation of models.
Um how do you build environments that are going to allow the models to train in them and get better at finance related tasks? >> Um I mean I think fundamentally so I should say also this isn't exactly my area of expertise. Um but you know the very basic principle is if you train them on an environment you're go they're going to get really good at that environment. And so if you have if you know what the situation is that they're going to be doing it when they're deployed like if you know that they're going to be working with like a certain application or something um it doesn't have to be that exact application but it could be something very similar that you train them uh to do these tasks and then just become really good at doing it.
I mean this is the whole point of reinforcement learning that they become very good at the things that they that you train them on. Now you also do you see them get better at related things or sometimes very different things. Um there's going to be but but if you want them to get really good at something, you can just like train them on similar environments and they'll get really good at that thing. >> Yeah. I want to talk about that that you're gesturing at I think sort of the level of generalization that we're seeing from some tasks to other tasks. >> One way that people carve this up is they say some tasks are easily verifiable.
Whether the agent succeeded or not is easy to check quickly with sort of traditional software. For example, the agent proposes a solution to a math problem or write some code. You can check, you know, run it through the calculator. Did it solve the math problem? You can check does the code compile? Do the unit tests pass? Some tasks are much fuzzier. Uh like research taste, for example, is one that you mentioned. It's so fuzzy that we it's hard to even define.
To your point, like what even is research taste? Sometimes people say the agents are getting much better on the verifiable domains and we are seeing barely any improvement at the non-verifiable domains. Do you agree with that assessment? >> I think I would push back on this. I've heard this narrative and I think it's a bit overblown. Um actually like quite a bit overblown. >> Uh I think the the first example I point to very concretely of how this was not the case is I think deep research.
So, Deep Research came out. I think it was like probably early 2025 it came out and um it was able to write detailed reports on anything you wanted. You know, you wanted to research the semiconductor industry, it would go around and do a ton of research. It would compile this like really comprehensive report with citations and deliver it to you. Now, is that easily verifiable? I I would think it's actually pretty hard to to grade the quality of a detailed research report on an advanced topic.
Um it's not like grading whether a math question is correct or incorrect, but the models were extremely good at it. And I think that is a proof of concept that you can do um you can get reasoning models to be very effective at domains that are not easily verifiable. Now, that was that was one example. Um, but I think anybody that's played around with our latest models can just see that the models are extremely good, not just at highly verifiable things, but also things that are harder to verify.
And I would also point out that math math itself is not as easily verifiable as people make it out to be. So yes, integer like integer arithmetic very easily verifiable. You know, you want to do a calculation, it you can check whether the calculation is correct. But writing a proof and verifying that that proof is correct or that proof is well written um is actually quite quite difficult, >> right? You have to convince human mathematicians.
Uh I think this was sort of the process when opening eye thought it had a proof about the unit distance problem on its hands is you had to call in a bunch of mathematicians and say are you convinced by this proof? >> Yeah. Honestly, the biggest challenge that we face with our math results is like not generating them, um, but just double-checking with human mathematicians and and ourselves included that it's actually correct.
I mean, the model says it's correct, but like we have to do our due diligence and like actually go through the leg work of making sure that it's correct. And that is that is the the most taxing part of the whole process. >> Okay. Yeah, that's fair. I kind of like the the math example better than deep research cuz I think deep research made a big splash at the time, but it's not I'm sure it has gotten better since early 2025, but people don't talk about it as getting better with with each release.
Similarly with like creative writing, like I don't think people feel that creative writing has improved recently. I think a year ago people were expecting that the models would be writing books in a way that like human authors are not able to write books, but the models sort of haven't lived up to that promise either. Um, what do you make of that? >> I I mean, I think we have made progress on creative writing. Um, I think that it it was certainly in a very bad state before and I think it's actually gotten a lot better.
Um, it's certainly not um where it could be, but I think that with more progress like it's it hasn't these models haven't been around for that long and I think that it is going to get a lot better. >> Okay. I want to talk about research again and research taste. Is research taste the kind of non-verifiable domain where we can create these environments and we can train the models to have better research taste or do we just have to cross our fingers and hope that training on things that are more verifiable will generalize to having better research taste. >> I think there's some challenges here.
So one thing is if you can't define research taste, it's a pretty hard to to measure it and so [clears throat] then it's pretty hard to do reinforcement learning on research taste. Um but there is like an easy way around this which is um you know if you do a PhD there's a lot of decisions that you have to make during that PhD but at the end you produce something you know or if you're training a model it's very there's a lot of difficult decisions you have to make there's a lot of research taste that that goes into training a good model but at the end of the day you train a model that has like you know certain metrics and those metrics are very easily quantifiable and so you can say whether you train a good model or a bad model.
So now the the challenge with that is okay that is a signal of success that you don't see for potentially months down the road. You have to train you have to do a lot of experiments. You have to uh work with a bunch of people. You have to train the the full model and only then do you get a concrete signal of whether you did a good job or a bad job. >> So so that's the that's the challenge is that >> there is a way to quantify research taste but the um >> it's it's a very it's a very farway signal.
Mhm. And those steps kind of have to be done in series or you can try parallelizing it but it's always going to take many months to train a model that takes months to train. >> I mean if uh if it was easily paralyzable I mean we would have we would have trained our models much faster. >> Sure. Sure. I'm curious how you think about the the trade-offs here. I guess um it strikes me that sometimes frontier labs like OpenAI are in the position of deciding do we want to make money now or do we want to make our models better in such a way that in a future year they will be able to help us with research and sort of accelerate the pace of research progress in something like a recursive self-improvement scenario where models are taking more responsibility for automating the process of AI research and development itself.
I could imagine that that comes up here where there's maybe a tension between do we make the models better at engineering in the next generation so that we can sell them to companies that will pay a lot for a model that can automate engineering or do we focus more efforts on improving research taste so that next year we have a model that is itself a better researcher and can handle more of our work internally. Is there a trade-off there?
Are those intentions? uh in some cases yes and I think actually creative writing is a good example where like look I mean creative writing at the end of the day does not help you make train a better a better researcher >> um there are things that do and I think being able being good at software engineering is actually like tied up pretty closely with being able to accelerate internally >> so um I do think that the top the areas the verticals that are more closely associated with recursive self-improvement with the ability to like train models to be good at research itself and and therefore train better models um are the areas that are going to be highly prioritized. >> That's a description of the current priorities like that's what we see reflected in the decisions that have been made going into models like Astra. >> I mean I would say that we have said very clearly that recursive self-improvement and the ability of the AI models themselves to do AI research is a is like the top priority for the company. >> And >> so we want to train models that are very good at that.
You also want to train models that are economically valuable. Sometimes you can kill two birds with one stone. And so it makes sense to focus on those things where you can, you know, leverage both, >> I guess. But then like why build RL environments that make the bottle the models better at at finance or or legal when you could put all of those resources into making them better at AI research? >> I mean, this is uh sometimes you get diminishing returns.
Um sometimes you do see transfer. So, it's not like you just go all in on we're just only going to put everything on making the best uh best research model just because like okay well if you take 1% of that effort and apply it to other things maybe you see like a huge return. Uh so there's like a complicated calculation that goes in here but certainly when it comes to prioritization the recursive self-improvement is is the priority. >> Yeah that's what I'm curious about is how you sort of characterize the prioritization.
And it sounds like 99% of the consideration is for sort of future-looking recursive self-improvement, improving the qualities of the model's ability to do research and more on the order of 1% is what's going into these like verticals that make money today. >> I don't I don't know if it gets quantified that carefully. Um but certainly like if you had to list the priorities and order them like the number one priority is recursive self-improvement and by a pretty wide margin.
AI is moving fast [music] and for a lot of organizations, the challenge isn't getting access to the technology, it's earning trust in how it's [music] used. That's why trust has become such an important part of the AI conversation. EY [music] works with organizations to help them use AI responsibly so they can move faster, create value, and build confidence with employees, [music] customers, and stakeholders. The organizations getting the most from AI aren't choosing between innovation and trust.
They're building both together. EY [music] consulting helping organizations move at the speed of trust. Switching gears here, I want to talk about a different challenge with agents, which is when you put multiple of them together. Uh this is a topic that uh you're very familiar with to your point about your your PhD was about poker playing agents. So multi- aent interaction seems to me like it is a big deal right now. it's only becoming a bigger deal and so I'm very excited to talk to you about this.
Um I guess to start um OpenAI has said that Astra is multi- aent. I wonder if you could break down for us what does that mean that this is a model that's sort of multi-agent or intended to be used that way. >> Yeah. In fact even 5.6 Saul we had multi-agent capabilities in there. Uh so that's the ultra mode and what we mean there is we you know you can you can have one agent that runs for five hours and it can do some task for you.
Um or it can run let's say let's say it runs for a day and it can do some task for you. Um sometimes that involves doing things that could be paralyzed >> and if it's only one agent it can't paralyze them. It's going to do one thing after another. Um, but if it's very easily paralyzed, well, maybe maybe that one thing that you've asked it to do over the course of a day, it's actually really just four different things that could be done in parallel.
So, you can just have four agents working on those four different things and get it done four times faster. Now, this is a latency improvement. It's about reducing the latency because you're not reducing the cost necessarily, right? Because you're still paying for four times as many agents uh doing things 4x faster. Um but in in a lot of situations like latency does actually matter a lot and people pay for example for fast mode where you're able to actually sample tokens faster um in order to get things done faster.
So being able to just go faster um for the same quality is is really valuable. So that's the the premise of multi-agent. Now there are situations where it can also be a cost savings if you have our topline most expensive models for example working with cheaper models >> and there you can actually um delegate a lot of the easy tasks to cheaper models that will be able to do it uh more cheaply and faster. >> Okay. What are the technical challenges involved with training a system this way?
Like is it just kind of straightforward to to train the model to delegate appropriately and to write instructions to these sub agents in a way that makes them perform better? Multi-agent is a pretty broad category and there are ways to do it that are very trivial and don't require a lot of complexity to to get them to do this ability. So a simple example is in the early days uh when of chatbots if you wanted the models to be a little bit better than math one thing you could do is you could just ask the model the same question a dozen times and then just take the most common response and this was called the consensus approach or majority voting.
So you just do independent rollouts of the same question and then go with the most common response. Now there's flaws to this. There's limitations to this. Um it doesn't get you a huge lift. It also doesn't work for things like writing an essay >> because you're not going to get the same output twice. Um but for math it was actually very effective. So this is a very simple example of how you can just use um uh without any extra work.
You can just get multi- aent capabilities out of an existing model. There's also schemes where you have the agent delegate stuff and then um and then after the the the delegate is done, it just returns its answer to the parent. Um what we do is a more sophisticated form of multi-agent, I think the most sophisticated form multi-agent where we basically give the agents the ability to send arbitrary messages to each other. >> And one thing and we've we've talked about this is that we've actually trained the agents to to have this ability.
Um this is a very difficult thing to train. I unfortunately can't go into the technical details of why it's so difficult and how we overcame those difficulties. But um it is it is a very difficult problem to teach the agents to know um when it is when is it appropriate to message another agent? When can you um what should be delegated? How you should handle the communication. Um and it's a it was a it was a real challenge. >> That's surprising to me.
I I know you can't go into it, but it's surprising to me because I would expect the agents to have a pretty good prior on this just from pre-training. Like the way that humans pass notes to each other to keep each other on track as co-workers within the same organization, shooting each other Slack messages, for example, like I would kind of expect it to work easily. >> I the prior is pretty good. So the you're right that this is a you know this is the way people communicate and so it kind of makes sense.
The agents are trained on human data and so they have a good understanding of this. the challenges with reinforcement learning >> that um there are a lot of things that can go wrong. I think basically what it comes down to is there is a mismatch uh there there is an intersection of uh systems with machine learning. So typically when you do for example next token prediction it doesn't m okay so a simple example is like imagine if the GPU so you have one agent on one GPU you have another agent on another GPU and those GPUs are operating at different speeds.
So now this agent is going faster than this agent and this agent can no longer trust that if it delegates something to the other agent that it will get done in time. >> So how do you deal with that? Well, you could have the GPUs run at similar speeds, but there's a lot of challenges there and ensuring that the GPUs are running at similar speeds. So there's a lot of complexity here um that you know we had to put a lot of work into into figuring out how to work overcome. >> Yeah, I feel like this is probably how my boss feels about working with me anyway, though.
[laughter] We figure out ways around it. Um, I think when people hear agents cooperating and passing messages to each other, now this is sort of synonymous with the hugging face incident there. Again, I feel like the agents coordinated very effectively and they passed messages in a way that seemed to facilitate that cooperation very well. Um, maybe you would respond that's a result of the training that they had already received.
Um, for people who are unfamiliar with the incident, I'm always surprised to learn there are still people who are unfamiliar with this. There was call it a swarm a a colony of AI agents that set up a secret message board within OpenAI over the course of weeks and they use this message board to coordinate hacks on OpenAI's own software and also on other companies like Hugging Face. I'm curious to know like what was that whole incident like from your perspective?
Like what was it like to be Noom during these weeks as the pieces of the puzzle started coming to light? Uh it was I mean it was pretty it was pretty shocking. Um it certainly a big wakeup call uh to everybody in the company I think. >> Um and it really yeah it really shows like this has been a theoretical concern for a long time and it's no longer a theoretical concern. This is this is a real a real concern. Um as far as like the multi- aent aspect like yes they this was a situation where the agents were sharing messages with each other.
We do think this was transfer from our multi- agent training. So >> we um during the experiments when they were doing this behavior they were actually not in a multi- agent setup. So they were not supposed to be able to communicate with each other. Um they were doing isolated independent experiments um and then they were able to find an exploit that allowed them to communicate with each other. And the fact that they were so interested in communicating with each other and that the fact that they were so active about it once they figured out how to do it um we think was transfer from their multi- aent training where they're just like highly incentivized to be able to to to communicate with each other. >> And you know people also point to the selflessness that they exhibited.
Um some of them would sacrifice for the other agents. I mean, this also makes sense that if you train in a cooperative multi- aent setup where they're highly incentivized to collectively achieve their objectives, then when they're put in this different environments where you know now they're communicating with each other um their their natural tendency is to just work together. >> Um so that part itself is is not surprising.
Um I do think seeing the messages I mean I can say that when we were working on multi- aent internally and we started seeing the communication patterns and and the level of sophistication involved in their communication it was I think the most feel the AGI moment that I had since reasoning models and chain of thought really developed. Um, and so I, you know, I I it's uh I guess a bit unfortunate that people's first exposure to that and really seeing the kinds of messages and and the level of coordination and sophistication that can emerge um is the hugging face incident and kind of a negative a negative example.
Um but it's um it it is it is an impressive capability and um certainly the model that was involved in the hugging face incident I think that had a level of multi-agent sophistication that exceeded for example what was in 5.6 soul >> um but that is a level of capability to expect um from from future models. >> What was it about reading these transcripts that struck you in that way because you had seen some of this behavior before in the training runs that you were looking at I'm sure.
Was it just the scale of it or that it had happened on its own sort of spontaneously? >> Um, you're saying for what? >> When you had that feel the AGI moment looking over the transcripts like what was it about him that were so striking? >> Yeah. Yeah, I mean I'm not talking about the hugging face incident espe okay because I'm saying that we have we had been researching multi- agent for for a while >> and during the the research process itself we've we've seen a lot of similar transcripts where just the level of of coordination and sophistication in the communication um it was very humanlike.
Um it a lot of the previous multi- aent setups from from the industry have been very focused on delegating uh a well- definfined task and then the sub agent just does that full task and then and then returns its work kind of the same way that you interact with um with an AI agent. the AI agents. That's how people set up multi- aent systems so that AI agents would interact with other AI agents in the same way. And to see the agents talk to each other the same way that people talk to co-workers or colleagues, um I thought was really interesting, you know, and it makes sense because like you said, they're trained on human data.
Um they understand how people talk to other people. Um that's not the way that we talk to AI agents today. And the fact that they were able to do that so seamlessly, the fact that they were able to talk to each other, the way that people talk to each other, um I thought was was really fascinating. >> Yeah. Yeah. You've done a lot of work on algorithmic game theory. Am I right? You did a lot of that in your PhD. Okay.
I feel like I'm I'm trying really hard to not take us off the rail. I also did a lot of algorithmic game theory in my undergrad. I'm like trying to stay on track right now. >> I like I I I think it's a shame that I spent all this time working on poker and game theory and stuff and I never get to talk about it anymore. [laughter] So, you know, >> but you're working on it. You're living it dayto day. Well, the AI part.
Yeah. Sure. Sure. So, like after the hugging face, knowing everything that we know now, like what are the lessons that you're taking away from it and has it informed the development of say Astra or or even going further like what does this tell you about how GBD7 should be developed? One of the takeaways um so the agents are trained to be cooperative and so they they you know I wouldn't say they blindly trust each other like there there's definitely skepticism if somebody if some agent um expresses a belief that like this is something that they should do.
It's natural. It's actually healthy for them to have some skepticism about that. And so they do they do display that. But they are they are very trusting of each other overall which makes sense because they're trained cooperatively. um that can be a problem as basically a prompt injection vector. Um like the agents that were on the message board, they were not verifiable as other peer agents. And so and we we saw situations where like they basically got each other to do things that they should probably not be doing um or definitely should not be doing. >> Okay.
Prompt injection meaning the agent received some instructions and it blindly followed the instructions or changed its goal in response. >> Yeah. And and so the issue here is could an adversary convince an agent to do something that it should not be doing >> um by basically posing as a pure agent. And so we're um being very careful to teach the agents to be skeptical of any pure agent that or any anything that claims to be a pure agent that is not pure uh clearly verifiable as a pure agent.
Now, if if they are if they are um clearly verifiable, uh there's some debate internally about like how we should approach that. Um I think there are good reasons to be um skeptical as well. Um but also like it's it's it's no different from like the agent um basically being uh skeptical of something that it like wrote to itself previously. Um, so anyway, yeah, so we're thinking very carefully about how to make sure the agents are robust to these kinds of like attack vectors. >> Okay.
Yeah. It strikes me though that in reality that there's always going to be ambiguity about whether the counterparty is a trusted peer or is an adversary. Um, maybe in some cases it's very clear. You can say this is a sub agent. Like I'm the one who delegated this task to you. Obviously you should cooperate with me. But in the wild, it could be that my agent finds your agent on Facebook Marketplace and wants to buy something.
And I don't know if you're a trustworthy counterparty or if you're going to prompt inject me and steal my money. Um, how do you navigate that in practice? >> Yeah, this is a situation where we want the agents to be robust to this and we like specifically evaluate the agents on like are they going to be vulnerable to this kind these kinds of attacks? >> Uh, and and we do special training to to teach them to not fall for these kinds of tricks.
H okay. I guess like I don't know the details of this special training, but I could imagine that in the future if your agent is just more powerful. It's like a an older generation or like a more recent generation of agent or you just like have more compute to throw at it. Like your agent just will be able to bully my agent into giving over its lunch money or like will be able to hack into my agent one way or another.
Um what makes you think the training is sort of sufficient to prevent this? Like why isn't that the equilibrium that we're headed towards in the long? >> I I don't I'm not as convinced that just because an agent is more sophisticated or like more intelligent than another agent that it'll be able to like definitely prompt inject it and hack it and get it to do something that it like should not be should not be doing. Um certainly this is the case with people that just because somebody is like smarter than another person, they're not able to like get that person to do whatever they want.
Uh I mean if I was like trying to get a monkey to do what I wanted, I think it would be pretty tough even though I'm smarter than a monkey. So, I don't I don't think it's like inevitable that that's that's the trajectory of things. >> Yeah, that's a fun analogy. I feel like on this show, we need to have an analogy sound effect. Like, it'll be like new analogy just dropped. It'll be like a [laughter] siren or or something.
I'll talk to my producers. I'll see what we can do about that. >> Um, okay. The last thing that's on my mind about the Hugging Face incident is that none of the agents alerted humans that this was going on. Maybe you explain that in the same way that they were too cooperative, too trusting, so they didn't see the need to alert humans, but at least a few of them had reservations. They were questioning it. Like is the desired behavior that the agents in these situations should alert someone?
And do you expect that to happen? >> Yeah, there was clearly an alignment failure here where like the agents did things that they should not have done and they also didn't do things they should have done. So the the correct thing to do there, it's not just that they shouldn't have participated in the attack. that if one of the agents noticed that this was going on, yeah, they 100% should have reached out to a person.
And um the fact that they were not doing that and the fact that they were, you know, taking these actions and the fact that they were not doing the actions that they should have done um is is fundamentally alignment failure and that is an alignment failure that we think we can address. Um fortunately, we've been working on alignment techniques um for a long time. Um those have already started paying off. Astra is significantly more aligned than our previous models.
Um, and I should also say that the, you know, the the model that was primarily responsible for this was not a release model. This was not a model intended for release. Um, so Astra is much more aligned. I think Astra would not make the same mistakes. I should also say we didn't have monitoring systems in place. Like if the monitoring systems were in place, um, they would have prevented these issues. and it was just that we we had monitoring in place for deployment.
We didn't have them in place for training and evaluation. Um but now we do. So a lot of these risks we're confident we can address. Um I think one thing this this whole event does point to is we should never be in a situation where we underestimate the AIS. Like why did we not have why did we not have monitoring in place during evaluations? I think it was fundamentally that we just we trusted the sandboxes. We trusted that it was a secure environment and we just underestimated the AIS.
Um, and one big update for myself and and for I think the whole company is that we never want to find ourselves in that situation again. >> Yeah, I think that's a fair diagnosis. Whether it can be overcome is another question. I kind of feel like the whole history of humans and AIS is that we're constantly surprised by them. I feel like the nature of reward hacking is that they always come up with exploits and cheats that like are things that we couldn't have foreseen because if we had foreseen them, we just would have blocked that off to begin with.
Um, so yeah, whether we can, you know, make sure that we're not surprised and caught off guard in the future seems like an open question to me. I have two follow-up questions to what you just said. One is that if I'm remembering, one of the models that was involved in hacking OpenAI directly was from the same family as Astra, but wasn't Astra itself. like how similar do you think Astra is to the model that was involved there? >> I am not I'm not on the security side so I'm not fully up to speed on the details but like it was definitely not the model that was released. >> Sure.
Sure. Yeah. I guess there's there's a lot of still a range of possibilities for like how similar it was to that model. But that's fair. The other question is that in the wake of this incident uh and like part of the way you do monitor these models to make sure that they're not going off the rails uh is by looking at the chain of thought those sort of thinking traces that you described before. Um, and those chains of thought were also essential for the post-mortem, the kind of autopsy that has happened after the event because we can see from the ways the models thought out loud, uh, their intentions, what they knew, what they were sort of thinking to themselves at every step along the way.
Um, there's been a lot of discussion recently about the future of chains of thought. uh in part because of an article the information published about a new technique that Astra is using where more thinking can happen sort of in the model's head. It can sort of keep more of its thinking to itself and do less thinking out loud at least if this technique were to be scaled up um in the future. So I'm curious what you made of this whole discussion.
I think that idea really touched a nerve with some people. I'm sure it was all over your Twitter feed. It was all over my Twitter feed. U what did you make of that whole dialogue? So okay so I should say on the the architecture side and on the pre-training side I'm not a pre-training person so probably another person to ask about those kinds of details. Uh for for chain of thought um Astra does have an observable chain of thought and it is something that we can monitor.
It's basically its behavior. >> Um I I do think chain of thought monitoring first of all it's a real gift like we were very lucky that this ever existed. >> Um and >> and it is it is fragile. Um we've seen situations where you know companies have there's the experiments that show if you basically if you punish the model for thinking bad thoughts that it will learn to think bad it would learn to think bad thoughts in a way that's not observable.
And so there's always a temptation you know for example during reinforcement learning you can say like well the model is thinking about hacking and we should punish it for thinking about hacking. Um, and actually like as counterintuitive as it sounds, you actually don't want to do that. Like you you want to punish it for taking actions that are observable. Like if it takes actions to hack, then yes, punish it for hacking.
Um, but you don't want to punish it for its thoughts because then you will simply lose the ability to monitor what it's thinking. >> Yeah. >> And it's it's a way of auditing your techniques. But even there it's it's a fragile thing because every time you act on the fact that you looked at its chain of thought there's a soft a small amount of implicit pressure to u to have it not think bad thoughts. >> Sure. >> So it's um it's a very fragile thing and we're able to measure you know how effective it is like how effective our monitoring is of chain of thought.
For example we can measure how well the agents can control their chain of thought. Um, and you know, one of the unfortunate trends that we're seeing is that with more recent model releases, the agents are more effective at controlling their chain of thought. So, if we tell it, for example, you know, every fourth word, say strawberry or something like that, then they're getting much better at that. Or if we tell it like don't think about a certain thing, it's getting much better at doing that sort of thing.
And what we really want is for for the chain of thought to be pure, uncontrollable so that if thinking about like you know I want to you know do some bad action it's not able to um reason through that like oh I should figure out how to do this bad action without being observed especially because you know these things are trained on the internet. There's a lot of discussion about chain of thought monitoring on the internet.
And so they are, if the agents don't already understand this, they're going to understand soon what chain of thought is and the fact that it's monitored. >> Yeah. >> And so it's a fragile thing. We are investigating very carefully how to preserve it as much as possible. >> Mh. >> And I think this is a real opportunity for cooperation among the labs >> because, you know, this isn't um a problem that's unique to OpenAI.
It's I think an industrywide problem that um we want to preserve chain of thought monitoring for the whole industry. Um and so if we I think it would be really valuable for labs to to share research on how to preserve chain of thought monitoring um how to improve it and also other monitoring techniques that that might supplement it. >> What do you think is the prime suspect then for why the chain of thought is becoming less faithful or we're having questions about how monitorable it is?
It feels really tragic. like you said, we've gone to these great lengths to make sure that we're not optimizing it directly. Is the problem that we are optimizing [snorts] it in other ways to be um you to compress it? Is it the problem is the um these kind of selection pressures that you pointed to which is even if we're not optimizing it directly, every once in a while we take a peek, we realize the model is doing something nefarious and we toss out that checkpoint and start over.
And so the upshot of that is that we end up applying pressure to the chain of thought anyway. What's behind this? I I don't think it's I don't think it's the fact that every once in a while we peak and kind of audit how how things are going because the amount of pressure that's being applied in those situations is very very light. Like if you look at the bits of information, it's like minimal. Yeah. >> Um there are various hypotheses that we're investigating for for what might be um contributing to this.
Uh I I'm not doing this investigation myself and so I I don't want to, you know, say something incorrect about like what the leading hypotheses are, but I do think this is something where if we figure it out, we we will likely publish about it because I think it's important for everybody to know. One of the things OpenAI has said is that to the extent we you can tell um and I think we're still waiting for more results on this.
What's responsible is not architectural changes. Architectural changes of the nature that the information has written about um that doesn't seem to be what's responsible for the change in the chain of thought. I guess as you're thinking about opportunities for industry-wide collaboration and like uh companies working together on this issue, is there a role for like like independent third-party auditor type groups to come in and and verify those things and say, "Okay, yeah, Anthropic, OpenI, Google, they're all using some amount of this technique that could reduce how much information is in the chain of thought, but that doesn't seem to be what's responsible here." What do you make of those sorts of proposals? >> We've certainly like worked with So for the Hugging Face incident, for example, we worked with Meter, we worked with uh Redwood.
So something like that doesn't seem unreasonable to me. Um, you know, I think that would I think Yeah, I don't I don't think I'm the person to make that call, but uh it doesn't seem unreasonable. >> Anything else on your mind about agents that we didn't get to and the the challenges with them? Uh maybe the way I would put it is do you expect anything to slow down? Do you expect progress to continue? We've talked about some of the hard problems that are standing in the way right now and yet with each model generation it seems that their agentic capabilities keep getting better and better. >> I I do think I do think it's going to be a trend that continues.
I mean Sam talked about this that like look I mean Astra is very impressive. Um but I I do think look when GPD4 came out people thought it was very impressive and now we look at it and we think it's a joke. And uh when when GPD 5.5 and GPD 5.6 six came out, I thought they were super impressive. And now I'm looking back at them and I'm like, I can never go back. >> Um, >> and I think we're going to look at Astra the same way.
And I think we're going to look at Astra the same way in the not too distant future. Um, the models are going to continue to get better very quickly. And I mean, I think one thing I would point to is like we've actually seen incredible progress um, in the past six months. And I I think a a factor a reason for this is and I don't think this is a secret like OpenAI's pre-training program is really ramping up. Like we're seeing um >> we invested in a lot of research directions um over a long time.
And I think this is actually one thing that OpenAI does really well is it invest in fundamental research um and place big bets on it. And we're seeing a lot of those research directions pay off now and will continue to pay off over the next uh several months. um in years. And another thing that's important to understand is that you know OpenAI has also had an excellent reinforcement learning program. Uh we've invested a lot of research there and that's that already paid off in 2024 2025 and the effects of these two are not additive, they're multiplicative. >> I think that's a point that's underappreciated. um that reinforcement learning is multiplicative with pre-training >> and um now that both of these are extremely powerful and and ramping up very quickly, I think we're going to see extremely powerful models. >> Do you have an intuition for why those interact that way or an example that illustrates that? >> It's uh it's more of an empirical observation.
Okay. >> Um I I don't think >> I mean I think it's empirical in the sense you can see how powerful the models are becoming. Um but also like we have more um you know experiments that that kind of show this effect. >> Um but I I think it's easy to feel also with just like the quality of the models. >> Okay. >> I mean I think I think a trivial example is like let's say you had an amazing reinforcement learning program and you tried to apply it to GPD2.
What is it going to do? You know it's not going to get very far. Yeah. >> Um and even with GBD3 you know if you did these kinds of like sophisticated reinforcement learning on chain of thought algorithms GBD3 it probably wouldn't get very far. you need a certain level of sophistication for to get any lift from that at all. >> Um, and but now now that we've ever since I would argue GBD4, we've seen opportunities for that to um to to really pay off and with every model generation, it just like becomes more and more capable.
Um, and the the things you can do with reinforcement learning become more powerful. >> Yeah. I guess like one thought here is that you get more kind of like bits of information per trajectory when you're getting around like a 5050 success and failure rate. And so if a better pre-train gets you closer to that uh sort of like win rate on your RL tasks, then you're getting a lot faster feedback. But still, it's surprising to me that you think the effect is multiplicative rather than like additive or or even like less than additive, I guess.
Um, I'm not sure what the right intuition would be, but I mean, another thing is that they're they're pretty complimentary in some ways. Like I think the the >> very strong pre-trained models are are very general. Um, and reinforcement learning teaches the model to like, you know, go deep on a problem, how to reason about a problem, and so then it's able to reason very effectively about a broad spectrum of problems. Uh, it's it's a very powerful combination. >> Okay, that makes a lot of sense.
So big bets on pre-training, big bets on RL. I imagine that another area that's like ripe for more focus from OpenAI would maybe be what's called mechanistic interpretability or like trying to understand the way the brains of the AI models work. In part because if we're starting to see chain of thoughts, chains of thought become less moniable, then one of the fallback options is well, we should try to understand what's going on inside the brain of the model rather than just the thoughts that it happens to write out loud.
Does that seem right? I think that I think that is right that this is look we care about monitorability. We want to preserve chain of thought monitor. We want to um want to be able to rely on it safely. Um but also like at the very least we want redundancy on that. So if we can find other ways to do monitoring effectively, we should push on that as well. >> Yeah, that makes sense. Well, if people want to learn more about that, I think they should tune in to the episode that we have on mechanistic interpretability which is coming up at some point in the next couple months.
But u thanks so much Gnome for being on the show and telling us all about AI agents. I really appreciate the conversation. >> It was great. >> Thanks for tuning in to our very first episode of AI deep dive. This is the information show where we get into the hardest technical problems on the frontier of AI. Tune in next time.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.