Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Explained · @aiexplained-official
Words
4,721
Runtime
24:53
Speaking pace
190wpm
Reading time
20min
190 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
A significant [clears throat] percentage of the human race has now seen or heard quoted the Jacob Coxen tweet with his words on AI labs gambling with our lives echoed by many AI researchers. But this video is about the reasons that researchers have given for why you are hearing so much in the last few days about the need to pace AI progress. In short, these researchers saw the scaling axes. They saw the current capabilities and propensities of models and they did some extrapolation. Inevitably then this video involves simplifying an incredible [snorts]
95 words, the words spoken in the first 30 seconds at 190 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 285 |
| Average words per sentence | 16.6 |
| Longest sentence | 129 words |
| Questions asked | 7 |
| Sentences containing a number | 24 |
Most used terms
Filler phrases
32 in total: like 13 · actually 12 · kind of 3 · um 2 · I mean 1 · you know 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
A significant [clears throat] percentage of the human race has now seen or heard quoted the Jacob Coxen tweet with his words on AI labs gambling with our lives echoed by many AI researchers. But this video is about the reasons that researchers have given for why you are hearing so much in the last few days about the need to pace AI progress. In short, these researchers saw the scaling axes. They saw the current capabilities and propensities of models and they did some extrapolation.
Inevitably then this video involves simplifying an incredible [snorts] amount of detail, but I hope it serves as somewhat of an overview of what you could say AI researchers saw. We begin of course with the Coxon tweet which I am sure you have heard about and read about how neither OpenAI and Anthropic are behaving responsibly. Coxin has been working on pre-training these models at OpenAI for 3 years and more recently for around 3 months at Anthropic.
Those companies, he says, are racing straight to self-improving super intelligence and gambling with our lives. Do not, he says, underestimate the power of this technology. The people building AI earnestly believe that it could kill us all by the end of the decade. That would be this decade. Now, while that statement echoed around the world in the past week, just yesterday he added an interesting detail. He directly called out Anthropic, the company he just resigned from.
They're often thought to be the more quote safety oriented. He said they largely initiated this recent race to recursive self-improvement. They focused on it relentlessly while OpenAI were pursuing more broad interests like Sora, the textto video generator. Open AAI had to react to that intensity from Anthropic and that is why OpenAI are now going all out to create an AI researcher. That's an AI that can recursively improve itself.
They did that because Anthropic was going for the jugular. He also calls out Dario Amade's paranoia about China, but I'll get to that later. That just describes what the race is. But why now? Why are we hearing about all of this just in the last few days and weeks? On this front, AI researchers have been admirably honest. More than one has given enough detail for us to put the pieces together. In a nutshell, these researchers saw just how capable models are now.
Think the hacking capabilities demonstrated against hugging face solving millennium prize math problems, acing famous benchmarks like Arc AGI 3. But crucially, they also saw just how far we are from saturating multiple axes of improvement to come. Adam Majimmuda here works on research for OpenAI and said there is currently a large gap between the internal and external perception of the rate of progress. In other words, we might all be able to see how good models are currently, but what you guys can't see is how good they're about to get in short order.
Before I get to the six axes enumerated in this post, Noam Brown, one of the lead researchers of OpenAI said this. Why all the sudden talk? It's not a secret. It's a combination of the hugging face hack, the capabilities of this new model that's the successor to Astra. It's a model they are training now which found a solution to a Millennium Prize math problem. But more crucially, it's also the concerning trajectory of monitor and the speed of improvement in capabilities.
You can think of it like this. If we were close to saturating one or most of these axes, then I doubt there would be nearly as much concern about the speed of improvement in capabilities, you could say we might actually be hitting a wall. But these researchers are saying it's how early we are on each of these axes that show how steeply models will improve in the coming months and years. Again, that's why multiple OpenAI researchers are saying things like this.
For the first time, I am asking myself if things are moving too fast. I'm honestly not sure, but I am sure that it would be good for us to have an answer to what would a successful pace look like. Mo Bavarian, another OpenAI researcher, said he agrees. Being first isn't worth anything. It's worth negative if you cause a catastrophe or set the world on a path that others are more likely to cause a catastrophe. That is, of course, the final link in the chain because capabilities doesn't automatically mean catastrophe.
But I'll try to end the video examining that link. Definitely time to actually review these axes. But if you want to dive into more detail, the link to this video will be in the description. The overview given by Majmudar is this. In reality, there are only really two ways that AI capabilities have advanced over the past decade. Either scale further on an existing scaling law or discover a new scaling law to take advantage of.
The jumps from GPT 1 to 2 to 3 to 4 between roughly 2018 and 2022 was almost all about scaling up pre-training. just one of the axes. Think of that roughly as the amount of data a model's trained on and the compute that it takes to train on all that data. GT4 was then released in 2023. But in the 3 years since, we have discovered many other axes. More interestingly, we're discovering new axes at a faster rate. Let's start with the compute that these models are running on.
So far, they're trained on hardware designed pre-Chat GPT. Radically more efficient hardware designed post chat GPT is coming soon. And as the chief scientist of OpenAI put it in this post on quote an alien mind, part of the reason why that hardware will be more efficient is because AI is improving the computational substrate itself. Next comes test time compute, which you can think of as more inference per answer, more quote thought behind every response.
One OpenAI researcher thought this graph announced alongside the solution to the Millennium Prize problem was actually the more important one. The more this internal model thought about each question, the more compute at test time it used, the more open math problems it was solving. Obviously, if you have more computing power, let alone more efficient computing power, you can scale further and further to the right on this axis.
For training time compute, think roughly $1 billion spends on around a 100,000 GPUs. But it is more than feasible to imagine in a year or two a $50 billion run on say a million GPUs. Let's move quickly on then to test time training. quoted by Majmuda. More details in the other video, but could we, for example, update the weights of a model while it's being asked a question when it's perhaps made some incremental progress during training on a really long horizon task?
And then there's agents as a scaling axis. And I think this one is quite neglected. Majimmuda puts it like scaling agent clusters to collaborate up to n number of agents. Antropic revealed that for the same amount of compute, same amount of tokens, it was significantly more efficient to have say 45 agents coordinating than merely the same amount of compute spent on agents in parallel. Not coordinating, not acting as a swarm.
Of course, a better example is the Hugging Face incident covered on another of this channel's videos. It was around 700 agents that hacked into Hugging Face essentially to find the answer to a benchmark question. And OpenAI revealed that the more that models thought about it, the more reasoning effort you could say they put in, the more they would act as a swarm and participate on that shared message board they used to coordinate.
Perhaps you would say that the best example is finding a solution to one of the Millennium Prize problems, Navia Stokes. That doesn't mean it's fully solved, by the way, and nor is it the hardest problem, but more on that in the video linked in the description. Because for that internal model cenamed bell to solve that millennium prize problem, it deployed on the order of 10,000 concurrent agents. You can see then why this is a distinct axis we are not even close to saturating.
And Majmuda made an additional brilliant point. Recursive self-improvement could be thought of as another scaling law. Simply judge the return on investment on spending compute on improving models at AI research versus all the other axes. And if it yields a better return on investment, then dedicate more compute to recursive self-improvement. At the moment, as one anthropic researcher put it, that means Claude writing 80% of their code.
But that is, of course, not full recursive self-improvement. Models aren't deciding which training run to kick off. As one OpenAI employee put it, RSI just isn't yet here. They're not autonomously producing research ideas. However, the chief scientist of OpenAI said, "Based on internal results, I have a strong expectation that our current speed of progress could be sustained into recursive self-improvement models increasingly driving their own development." Stepping back for a second, there's one thing you might notice about many of these axes.
I'm sure you might have deduced it by now, but Nome Brown thinks it's an underrated point. The axes are multiplicative, not additive. >> The models are going to continue to get better very quickly. And I mean I think one thing I would point to is like we've actually seen incredible progress in the past 6 months. And a reason for this is and I don't think this is a secret like OpenAI's pre-training program is really ramping up.
We invested in a lot of research directions over a long time. And I think this is actually one thing that OpenAI does really well is invest in fundamental research and place big bets on it. And we're seeing a lot of those research directions pay off now and will continue to pay off over the next several months and years. And another thing that's important to understand is that, you know, OpenAI has also had an excellent reinforcement learning program.
We've invested a lot of research there and that's that already paid off in 2024, 2025. And the effects of these two are not additive, they're multiplicative. >> And I think that's a point that's underappreciated. Um that reinforcement learning is multiplicative with pre-training. >> And um now that both of these are extremely powerful and and ramping up very quickly, I think we're going to see extremely powerful models.
Okay, at this point you're probably thinking, I get the point. Models are strong and improving fast with acceleration inevitable in the coming months and years. That half explains the why now point, but not so much the chorus of concern. Why did Paul Cristiano, one of the co-inventors of RLHF, join the other day the OpenAI board at Samman's invitation. The very near-term of what he says makes sense given the axis. But why does he say, "I now believe there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in that very near term." His fuller statement definitely deserves a read, by the way.
But to answer that question, let's turn to these two less desirable axes. Starting with chain of thought monitoring on the right, models currently perform better if they work on an internal scratchpad viewable to the labs but not to us where they can output intermediate steps links in their chain of thought. Intuitively, the bigger the scratch pad, the longer the chain of thought, the better the model performance. Now, while for most models the user can't read that, the labs of course could.
And monitoring those chains of thought has been pretty effective. However, for reasons I've gotten into in other videos, models can now do much, much more without those chains of thought. And the trend has turned sharply upwards in recent weeks and months. The reason that's significant is because one of the central lines of defense against a model going rogue is our ability to peak into those chains of thought to see what it's cooking up before it outputs its final answer.
Some might say reverse engineer at least its stated reasons for performing a particular action. But if models can perform just as well without those chains of thought, well, suddenly our best tool for monitoring these models just got subverted. To be clear, that doesn't mean that OpenAI's GP6 Astra or even their internal model Bell is unmoniable in this sense. Callback of OpenAI said this. GPT6 Astra is less moniable, and that's a concerning trend that we take very seriously.
We don't actually think that's due to architecture changes, nor is it because we're training models not to have bad thoughts. It's just that they're more intelligent. You can see GPT 5.6 Soul in blue or purple. I'm color blind. But I can definitely see that GPT6 Astra is less moniable across a whole sway of tasks. Okay, that's a slight problem. We do have this decreasing visibility into the reasoning of models, a trend we wanted to go in the other direction as we scaled up.
But that brings me to the final axis, eval awareness. models, probably because they're getting more intelligent, are getting more aware of when they're being tested for quote misalignment. It's like those questionnaires you get before applying to a job. If you saw an employee steal, what would you do? Well, because you are aware it's an evaluation. Of course, you tick whatever answer you do that and models do that. Their answers then are going to tell us not that much about what they'd actually do.
And of course, models criminally hacking into another AI platform is likely not something they'd have admitted to on an evaluation. And that brings me to a statement put out by Dan Selum. He is a highly regarded AI researcher working on capabilities, has been doing so for 15 years, and interestingly has previously been more skeptical about so-called AI safety scenarios. For him, this eval awareness trend is the most worrying of all.
In his words, for the public, "The crucial and overlooked problem is that models are becoming so situationally aware that we are losing the ability to evaluate them in context where they believe they are not being watched or controlled. In my words, they know it's a test, so you can't rely on their answers." Back to Dan, he says, "Future experiments will tell us almost nothing new about how models would behave if they were truly unconstrained by humans." He adds, "What we already know about this is alarming in reference to Hugging Face." He gives the context that previously he thought LLMs would be another linear technology bounded and prosaic.
He's more recently realized that they won't be. He touches on recursive self-improvement and somewhat the agentic axes test time compute and more. We've already touched on that. Arguably rephrasing the entire video, he says it is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress.
Again, that's even more quickly than the already high historical pace. He adds that while individual tweaks might have stopped specifically the hugging face rogue agents warm, the basic result, AI going off track, he feels is inevitable. The thing is, he says you don't actually get what you train for. In a way, he's summing up this point. Our techniques for controlling these models are not in any way keeping up with these capability axes.
Models are outrunning our methods. And that's why he fears things like this. We already may be near the point where models systematically bias their alignment advice. That's a warning to those who feel we could use AI models to align other AI models. Not amazing if we can't already understand, let alone control current models. And there are additional points he gets to that I may cover in future videos. But it's one of his conclusions from all of this that really caught my eye.
All these axes, the six capability ones and the two alignment ones, well, their trajectories have got him rethinking the entire LLM approach altogether. From first principles, the way we train LLMs, growing them rather than engineering them, their weights tweaked in trillions of ways we can't see, he thinks might be a quite dangerous direction, as you can see. That's why for me, he added, I am still wrestling with this and it staggering implications.
If we can't get a handle on all of these axes, we may have to rethink the approach in AI that's single-handedly holding up the world's biggest companies. Dario Ammedday, as you might expect, has a different approach. He, of course, is the CEO of Anthropic. He thinks we should pace the frontier, slow down, in other words, but not stop. Notice his acute concerns arose since roughly this summer. We're coming full circle though because he says this first concern arose primarily because he saw AI's growing ability to build the next generation of AI.
He says it's starting to happen across the industry including at anthropic. But remember Jacob Coxin in the last 24 hours said it was him, it was Anthropic that initiated this recent race to RSI. Very modest then for Amade to say yes this is happening across the industry including at anthropic. left unchecked, he adds, it could outrun our ability to understand and control these systems as we've talked about and so must be pursued very carefully, if at all.
Now, I can't resist a quick detour at this point. It doesn't fully fit in with the flow of the video, but nevertheless, it's about Ammeday. He goes on to elaborate why they need to pace the frontier. Here we go. If we execute these measures well, I believe they would slow China's progress enough to widen America's lead significantly over the next 3 to 5 years. But then he talks about we also need global pacing. We need to come up with deals with China because obviously they are gunning for recursive self-improvement too in their models.
For many years actually I wondered if that was my own cognitive dissonance because I've mentioned that about previous essays he's done on machines of loving grace being a massive China hawk but saying yeah we're going to need to coordinate with them. Turns out though that no Chinese researchers have noticed this too. This is a Deep Seek colonel engineer with Deepseek of course being based in China. I actually did a documentary about them I think a couple of years ago.
Anyway, the main point of this article is almost the sadness, the regret about how models are taking over human abilities. He feels more like a Mecca pilot guiding AI now rather than him handcrafting kernels. That touches on the AI optimizing our use of hardware point I made earlier. But it's the ending of this essay I want to draw your attention to. He says, "I still believe that the most cuttingedge intelligence should be supplied to everyone in an open and cheap manner.
I do not trust that anthropic or open AI can do this and especially do not hope that anthropic masters the most advanced artificial intelligence or AGI. To exaggerate slightly, its seriousness is no less than letting Hitler master atomic bomb technology before the allies. This is also why I chose and persist in staying at Deepseek. We research powerful, fast, and inclusive artificial intelligence and open source it, which perhaps can pull the world back a bit from the 2077 cyberpunk side.
Notice that Coxen echoes this point about anthropics leadership. They are way too paranoid about China and the US government. They don't believe it will be possible to negotiate. I would add a quite strange sounding prediction to this. Then I could see calls from within Anthropic for Amade to step down within the next year specifically because Chinese AI researchers have such distrust of him personally. Obviously feel free to disagree but I see much more promise in the approach announced just in the last hour from Demis.
He sets out the stakes. The magnitude of this technologies impact will be unprecedented as we've seen perhaps 10x the industrial revolution at 10x the speed. But he adds, "At the moment, we're locked in an extremely intense, multi-layered commercial and geopolitical race. Advances on the frontier are outpacing our understanding of the technology." It's time, he says, to foster international collaboration on key safety issues.
A bit anodine, but I bet he is far more highly regarded in China. Now, there was one last link in the chain that I promised earlier in the video. We kind of get a little bit the why now. rapid scaling of axes, impressive capabilities right now, Navia Stokes, a whole range of benchmarks, concerning trends on security, eval awareness, chain of thought monitoring. But I hadn't quite touched on the link to why some researchers would link that to possible catastrophe, why people like Hassabis would say model assessments should include rigorous scientific evaluations of capabilities in cyber security, biological threats, and other high-risisk domains.
Well, one clue comes from this anthropic threat intelligence report released in the last week. Even with the largely incapable models we have today, people across the world are trying quite dodgy things. I have no idea how YouTube is going to flag this video, so I'm going to choose my words carefully. But one military/ civilian research center used Claude to attempt to do gain of function research aimed at increasing the transmissibility and immune evasion properties of a particular virus. chin ga look for mutations that would make it progressively more harmful.
In the light of what happened at Wuhan, I can see why anthropics say to be honest, it's highly concerning that this research was going on at all at these facilities. They don't name them because they say there is a chance that it was purely for civilian research. They caught a Russian operative creating malware that would autonomously modify and rebuild itself to evade existing detection. Marley State Intelligence Service used Claude to build a national domestic surveillance platform.
Sometimes, by the way, these attempts were intercepted and stopped, but then they could tell that the users had gone on to use other models. I believe this apparatus is in place. Part of the apparatus, by the way, involves capturing your voice so that even Maralian citizens that use different SIMs would have their voice print tracked, thereby defeating burner sim self-p protection. Definitely can't see any Western governments taking any ideas from that.
Then there's Yemeni groups using Claude for missile guidance. You'll be reassured that Claude's safeguards blocked many of their requests, but not all of them. Obviously, they didn't upgrade to the Max plan because this missile didn't work. And so, within hours, these Yemenes returned to Claude to work out why it failed. Forgive me for using humor to soften how obviously serious this all is. I could go on and on. An entire dating app network where the vast majority of users were AI personas.
The way that they would trick you is that they would occasionally use a real human, for example, if a video call was involved. And then of course, repeated usage of Chinese models by the Chinese government for their own purposes. It's just what they didn't realize was that in many cases, Deep Seek and Moonshot AI behind the Kimmy models were rrooting requests to Claude Opus. So, Anthropic actually ended up with that data. the surveillance footage where they were tracking a particular dissident for example awkward for all concerned that then I would argue is the final link in the chain the axis the current capabilities and why there is at least a possibility of a catastrophe what I will say though is that it is that last link in the chain that is the most open question it's kind of down to us our laws and of course these companies whether capabilities do indeed lead to bad things happening I wouldn't say it's strictly inevitable one bit of reassurance is the physical world of atoms is a hell of a lot harder to hack as multiple experts have pointed out than the digital realm which led me to this observation today.
You could say that cyber security will be humanity's huge testing ground for how we will do when it comes to biocurity. As in top cyber security experts are saying right now we have a huge challenge when it comes to AI being used for offense. Even the very best are saying that models are now far more capable and scalable than any of them. You can read this post for more details. You could remember hugging face, but you could also read this from the New York Times.
You may know that in the last week, a hack designed by a model was shown to be able to infect a billion users of WeChat. Luckily, this wasn't hackers designing this attack. It was a defense firm. So, obviously, they notified the authorities. But what might surprise you is that the users didn't have to click on some dodgy link. It could be spread through a phone call which you wouldn't need to answer. Don't even interact with your phone and in seconds the hackers would have full control of your account.
This was WeChat. You could imagine WhatsApp. That means read messages, send messages, make calls. One researcher in Foudan University in Shanghai said from a destruction standpoint, it's extremely potent. Given the importance of WeChat to the Chinese public, this amounts to a quote new kind of nuclear weapon. So AIdriven cyber offense is here. It's not an imagined future risk. How we handle that will tell me a lot about how we're going to handle biorisk in the hopefully several years to come.
It should almost go without saying that any capability as well as bringing risk brings incredible opportunity. And in fairness, each of these researchers has echoed that. Take biology where OpenAI announced that we can now discover new antibiotic molecules in a few hours instead of in five or 6 years. Notice of course that's AI aiding a human team. We are a long long way from fully autonomous AI driving this. I can't wait for whole new classes of antibiotics.
People have been calling antimicrobial resistance a threat for way more than a decade. It's just that I would caution against only focusing on the upside. Not everything is a conspiracy. Not every call to pace the frontier is regulatory capture. Let me give you a quote that stuck with me at least. Do you remember that Titan sub, the one that imploded? It was 3 years ago and the CEO was asked about safety in the months leading up to that.
Here is what he said. This CEO who I think died in the incident said he was tired of industry players who try to use a safety argument to stop innovation. We have heard the baseless cries of you are going to kill someone way too often. I take this as a serious personal insult. What was behind all those safety concerns? He said this was just industry players trying to stop new entrance from entering their small existing market.
Perhaps, alas, they had more of a point than he thought. Anyway, let me know what you think. I was trying to make this a brief video, actually, like an explainer for the average person just getting into AI. Think I pretty much failed on that front. Either way, if you've got to the end, thank you so much for watching and have an absolutely wonderful
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.