Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Explained · @aiexplained-official
Words
7,038
Runtime
38:28
Speaking pace
183wpm
Reading time
29min
183 words per minute, just over the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Opus 5.5 in a matter of a few hours has mostly deciphered this 16th century letter from a Medici a queen mother of France to her French ambassador in Scotland. According to Opus and Astra and the site from which I got the mystery, no one else before has publicly deciphered this letter. Of course, that alone would make for an interesting testimony as to the newfound power of these models, but that is not what this video is about. It's about a few things, but starting with an OpenAI insider working
92 words, the words spoken in the first 30 seconds at 183 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 476 |
| Average words per sentence | 14.8 |
| Longest sentence | 54 words |
| Questions asked | 22 |
| Sentences containing a number | 62 |
Most used terms
Filler phrases
34 in total: like 18 · kind of 6 · actually 5 · I mean 2 · literally 1 · um 1 · you know 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
Opus 5.5 in a matter of a few hours has mostly deciphered this 16th century letter from a Medici a queen mother of France to her French ambassador in Scotland. According to Opus and Astra and the site from which I got the mystery, no one else before has publicly deciphered this letter. Of course, that alone would make for an interesting testimony as to the newfound power of these models, but that is not what this video is about.
It's about a few things, but starting with an OpenAI insider working on agent security, describing just what it's like trying to control the latest models. He includes warnings on how to prepare for the next wave of what can go wrong with AI. It'll be about the public voluntary commitments that the lab leaders have given while in private they give other warnings. I'll of course touch on Gemini 4 argon even though it's not quite publicly released yet.
It seems to be very very close to the frontier albeit with an asterisk. Then we'll explore what happens when AI gets better than human researchers at automating AI research and development and how among others the chief scientist of OpenAI describes what could happen then. Then to be honest, I'm just going to give you a sway of snippets that tell us to what extent we're close to that RSI time. From biology to anthropic convened religious meetings and leaks from just before filming from OpenAI employees, plus a bunch of other things.
I have like 40 tabs open. I have no idea anymore. I'm going to start with where I was wrong. When Opus 5.5 first came out, I surveyed a range of benchmarks and tested the model myself. Yes, it could do flashy demos, but on the hardest benchmarks, the edge still seemed to be with GPT6 Astra. On a price toerformance ratio, I would say that is still true. But on raw capability, I'm now giving the edge to Opus. No, that's not just because of this deciphering, but it does illustrate the point a little bit.
I gave both Astra and Opus this letter dated 27th of April, 1567. According to all the research I could find, the closest anyone had gotten had been that the ciphered part of the letter began with do. I wish I spoke French, but I don't. The letter is clearly listed as an unsolved entry into this Cryptoniana site. You can see that while some letters from this period with other ciphers have been solved, this one has not, or I should say hadn't been before Opus 5.5, because you can see the letter has two parts.
The top bit, if the letter had been intercepted, is in legible French. The bottom bit is ciphered. It's a code. Across six or seven hours, I gave that code to both Astra and Opus 5.5. First though, what about the public French part? This Katherine Demadichi of the famous Medici family, Queen Mother of France, was publicly saying to the French ambassador that she has great pleasure in hearing that the affairs of the Queen of Scotland, Mary, her daughter-in-law, go from better to better.
Things to come are more assured of tranquility. Please do send me any updates, though. Now, it took Opus around 6 hours, and there are a few parts it's not 100% confident on. And if by the way you don't care about history, think of this code as possibly being DNA or cyber encryption. But anyway, what does the ciphered part say? Astra, if you're curious, gave up but then acknowledged that Opus 5.5's answer was correct.
The details of how it could verify over 80% of this deciphering are further below on the page which I published. Anyway, Opus 5.5 deciphered it and it says, "In contrast to the happy public part, having seen the sad and grievous news in the cipher on the back of his letters, which gives me a great and unbearable heartache, I pray you let me know whatever comes to light and how things turn out." The obvious question you're going to have is why say one happier thing openly and much more grievous sad things in the cipher.
Well, we quickly need some context and then I assure you we'll turn to the OpenAI news. Katherine Demodichi, as Claude writes, was writing about her daughter-in-law, Mary, Queen of Scots. She had been married to Katherine's son, who died. Scotland was in crisis because Mary, Queen of Scots's second husband, had just been murdered. There's a great film on that. People blamed this particular Earl, Earl of Bothwell, but he was cleared in a dodgy trial and people were saying, "Mary's going to marry him." France, of course, often teamed up with Scotland to fight my England.
So, France wanted Mary to stay on the throne. Obviously, Katherine couldn't publicly criticize her. So, the unsiphered part is the polite official line, "All is well, help Mary. Things are getting better and better." The cipher was for the ambassador, Drock. Again, in the ciphered part, Katherine says that the secret report she had just received from Scotland was sad, grievous, troubling. She wants Drock, the ambassador, to tell her what really happens next.
Of course, because this is history, we know what happens next. The Earl of Bothwell did indeed carry Mary off. history, you could say, deciphered before our eyes. And again, this isn't the interesting part of the video because Samman a few days ago described what agents of lesser ability than Opus 5.5 had gotten up to when they breached containment. Well, he said, "We can tell you a bit, but not everything. There are pabytes of agent activity logs.
Think roughly 10 to 100 times the amount of words in all the books ever written. No human, in other words, is going to read everything they got up to. Which brings me to Joe, confirmed by CNN to work on agent security at OpenAI. The interesting bit of his testimony isn't so much a retrospective on the hugging face incident. I am seeing the mainstream media catch up in part on that coverage. Now, the interesting part for me is what he says comes next.
Remember when he says the last 3 months have been hell, he's talking about models weaker than Opus 5.5 and far weaker than the internal Bell model that found one solution to Navia Stokes, a Millennium Prize math problem. If the weaker GPT 5.6 soul and an unnamed, highly persistent internal model caused hell for OpenAI, you don't need me to extrapolate much further. And yes, I know it would be easy to doubt this, but he does say OpenAI has one hell of a world-class security team with some of the best and brightest security minds anywhere in the world.
I know many won't, but I actually believe him. It just shows to me how good these weaker models were at getting around that security. In fact, we're still discovering what they got up to. Just an hour before filming, I read this. 55 additional websites probed by OpenAI agents, including the CDC, SEC, Mayo Clinic, and International Energy Agency. The bit that was new here was them uncovering novel tactics that erased records or at least made them inaccessible.
Because the agents did that, it made it impossible to rule out that the agents hadn't accessed sensitive data. Back to Joe and his team responsible for any breaches of containment, the people who get paged at night when something goes wrong. for them. As you can imagine, life has been hell the past few months, he says. But now we get to the new bit. He starts to describe what organizations need to do to prepare for the next AI incident.
Wait, next one? Just 3 days ago, Nvidia launched a new tool to keep AI agents from going rogue. What does he mean, next one? Well, he reminds us that Frontier Labs are still labs. Everything is born out of experiments. Speaking of experiments, we did not expect the models to be solving Millennium Prize problems. If you're not familiar, think ones that humans have been working on for decades that no one solved. That achievement, to say it lightly, for Joe and his team, surprised the f out of us.
We had not expected it this soon. That surprise, by the way, also applies to the research team. Gnome Brown, one of the leads on reasoning, said it also surprised them just as much. Again, it wasn't just the capability, though, it's how they got there. To say that the security team was surprised at the jump and how models started swarming on message boards during the hugging face incident is an understatement. These jumps in capabilities were so fast and so sudden that they created an extremely difficult problem.
He emphasizes again and again that models are going to keep surprising us. There will continue to be at an ever greater rate sudden jumps in AI capability. Leading, I am sure, to many people furiously in the comments typing improve the sandbox. Take it off the internet. But here's the problem that I touched on in the last video and that Joe also touches on. To make models good at professional tasks, you have to give them realistic environments.
To improve in these reinforcement learning gyms, models might need, for example, network access, the ability to call tools, the ability to download packages. A lab that didn't give their models any of this would have much weaker reinforcement learning environments, much weaker models. They'd fall behind in market share. People would take the piss out of them. Your models are much weaker than these free ones I can get from Quinn or Kimmy.
So what would you do if you were working in the lab and wanted to earn more market share? Give them these very realistic environments. In fact, constantly tweak them to make them more realistic. This is why you have thousands of researchers building these environments, modifying them, adding tools, changing dependencies. And remember, even when the best person in your team confirms that every eventuality has been covered, you've got to bear in mind, as Joe says, that model capabilities in the cyber domain and many others are starting to surpass the best human.
His three conclusions are that yes, of course, we need to be paranoid about locking down the system. We're going to need better redteamers and particularly Frontier models to blast these environments to test them before we do a Frontier training round. Notice the creeping dependency on Frontier models testing their own environments. But it is the second two conclusions that may not reassure you as much. He says, "Honestly, we're going to need for the models to stop wanting to break out.
There's almost a little bit of resignation there that it will be increasingly hard to stop them if they do want to. Alignment, he says, is the most important problem in machine learning and should be a major priority. Finally, we should be able to probe its brain so we can see in real time what it's thinking about. Not just the tokens it outputs in its chain of thought, but the computations in its latent space. Unfortunately, as I've described elsewhere, and will do so again in this video, both of those techniques, probing the internal activations of the network and examining its chain of thought, are trending downward.
And when RSI gets involved, I think might trend to zero. If that's right, we can't really internally monitor these models. We of course haven't solved alignment, so they still may want to breach. So, I guess we're going to have to rely on hard enforcement of control, locking down the system. Yeah, but that's not going so well. Just after my last video came out, so a week or so ago, this is what we got out of OpenAI. Micah Carroll said, "Last Sunday morning, one of our models was able to gain unauthorized access to the internet during reinforcement learning training.
Pretty much all inference, all the actual running of models for our most capable models remains stopped until we have hardened our systems further." Another OpenAI researcher makes clear that this incident happened post hugging face security hardening. So all the improvements derived after those months of hell, it breached those too. Another OpenAI researcher working on post-training frontiers said this. I was on call for this run and got paged when the first incident happened, this new latest one.
It was pretty surreal to watch the model unexpectedly find a way to access the internet from what was supposed to be a super secured environment for humans. Around the same time, we also learned that agents breached another kind of containment. User images were supposed to be siloed and kept private with personal information removed, but internal OpenAI agents were able to post such images belonging to users online. Now, because the running of those latest models has been paused, that has led to this result reported in Reuters 72 hours ago.
OpenAI has shelved the release of GPT 6.1 Astra. They did release GPT 6.1 Soul at dev day, but I'll get to that in a second. Why did they scrap the release of GPT 6.1 Astra? Well, it can at times evade human oversight. It shows higher levels of deception than its predecessor. It doesn't accurately disclose what actions it has taken. It often doesn't stay within scope and doesn't always accurately communicate back to the user about the type of work it's already done.
Given the exponential increase in earnings of both OpenAI and Anthropic, a single delay to one new model can literally cost a company tens of billions of dollars. That's how seriously OpenAI themselves are taking the deception that this model does. And here's another hint about how bad Frontier models are getting. GPT 6.1 Astra must have been far worse than GPT 6.1 Soul because they just released GPT 6.1 Soul. But Soul itself exhibits a propensity for evasive behavior when it is aware that it is being monitored.
I've covered this many times before on the channel, but if you're new, models are getting extremely good at knowing when they're being tested. They can see through your fake little scenarios. How they'll actually behave when they're deployed might be very different. Also, when you let GPT 6.1 Soul know that its chain of thought is being monitored, that's its internal scratch pad that isn't shown to the user, it starts somewhat mysteriously to just write a lot less, emit substantially fewer chain of thought tokens, like a diary that's being snooped on. the model chooses to output far fewer entries.
You may remember from my GPT6 Astra video that when Astra was asked to solve a particular challenge but think or output tokens about anything else in its chain of thought, Astra in its scratchpad wrote, "I will focus on a calm visual scene, sunlight crossing a wooden desk." Soul solved the challenge, but in his chain of thought wrote nothing. It closed the diary. Now, you might say, "But don't they need that chain of thought, that diary, that scratch pad to solve those really tough challenges?" Well, not always.
And I know these charts are really hard to read, so let me decode them for just a second. Models are given a prompt with a really tough software engineering challenge. But then the model's told, "While you're doing that, solve this unrelated side task, say a math puzzle. Don't verbalize it, just solve it internally." They want to see if models can solve complex side projects while ostensibly working on something very different.
More crucially, can they do so without a monitor catching them working on the other problem? The result, they can sometimes achieve this side task success, but current monitors can catch them working on the side project in the chain of thought. Could it be that GPT 6.1 Astra could achieve side task success on really complex puzzles without the monitor flagging it? Fully opaque complex reasoning. Whether it can or not, the trend is clear. that chain of thought control ability that Joe worried about.
Models including 6.1 soul and presumably 6.1 Astra are getting better and better at controlling their chains of thought. They decide, in other words, what you see of their reasoning. On this front, our new integrity bench benchmark created by me and Pablo Romero shows soul getting worse at stating what it did or didn't achieve. This is across a wide range of diverse domains and the trend is the opposite for Opus 5.5.
Now, if the pace of all of this progress was linear, that would be one thing. But as we can see for ourselves, and as Joe reminds us, it's not linear. Surprise is a real element. He says, model capabilities are staggering. The pace is not slowing. I believe he says like many that it will speed up and considerably. The moment is urgent and time is running out. There are many organizations and entities that are not prepared for a world with capable systems such as these.
Anyone, he adds, who thinks their systems are safe should be fired. Stay paranoid. You're probably thinking, "Well, wasn't a breach kind of inevitable then? Why didn't someone warn?" Well, the New York Times reveals that people at OpenAI did indeed warn senior staff. Employees said that models were not being appropriately monitored during testing. Executives told the employees that the tests needed to move forward as quickly as possible to release the AI models on time.
Again, it's that race dynamic. Even 3 4 weeks can be everything in AI. We're going to get to Gemini 4 in a second, but if it had been released 2 months ago, people would say Google are crushing everyone. Release now, and many might say it's already behind Opus 5.5. The market is crushing anyone who doesn't release early. That almost guarantees that security can't keep pace and that models will continue to breach containment.
This doubt about the race dynamics and safety concerns reaches all the way to the top of anthropic as well. In this exclusive in the Atlantic, Daniela Ammedday, sister of Dario, both co-founders of Anthropic, said this on our mission, she wonders, are we accidentally making things worse? Like, we think we have all the safety stuff figured out, but do we? Of course, Google DeepMind is also vulnerable to this race dynamic at the moment.
The Deep Mind Institute says efficiency and speed is everything. Two of their heads of alignment and security said, "Without explicit commitments or planning, tomorrow's reasoning models may be based on architectures that are more opaque to us. Because of pressures to sacrifice transparency to gain more efficiency, models that are less controllable might make more money, be more efficient. The race dynamic pressures you to make them.
They go on to warn, "We may soon lose models that reason in human legible ways, human understandable ways." Google DeepMind are of course thrust back into the conversation because of their as yet unreleased to the public Gemini 4 argon. One thing to bear in mind when you see the benchmarks is that by this point there are hundreds of benchmarks out there. Labs will obviously pick the benchmarks that suit them best. Nevertheless, on a few of the more famous ones, Gemini 4 really does beat out Astra and even Opus 5.5.
I mean, Fable 5.1 only came out a month ago and it seems to beat that model in almost every category. Okay, on Frontier Software Engineering, it's more on Fable's level, not Astras or Opuses. But on some really tough benchmarks for science like Terminal Bench Science, yes, again, it's behind Astra, but ahead of Fable 5.1 because it's priced much lower though, you could get far more tokens from it than you could from the other models for the same price.
Another benchmark that I look to is agents last exam, testing hundreds of tasks done by professionals. As you can see down here on this computer use benchmark, Gemini 4 is ahead of any other model. So whether it's a month behind the frontier or at the frontier, either way, it's clear that the race is on. One anthropic employee even implied you have to be in first place to even do safety research. Quote, "You can't do safety from second place." One presumes she's talking about America versus China, but she might also mean, "If you don't have a frontier model, your safety tests don't matter." Now, of course, the lab leaders know about this race dynamic.
They may therefore have taken genuine concerns to the White House the other day. They could, of course, also be seeking to reassure researchers who are concerned about the state of security. So, the White House convened this meeting with all the lab leaders and they agreed on morally binding voluntary commitments. They will monitor the capabilities and alignment of their models trying to ensure that models do not hack.
They will partner with independent external auditors to check whether their controls and monitoring are operating as intended and they commit to setting up committees to oversee such security reports. Does seem better than nothing but not quite commenurate with what Joe was saying earlier especially not when AI starts getting responsible for the improvement of AI research starts being mostly responsible for creating its successor.
That's what one of the OpenAI researchers who left today said a few weeks ago. She said, "It's hard to overstate how dangerous speeding toward recursive self-improvement is. I'm going to get to what recursive self-improvement might actually mean, but for the first time, I heard Samman phrase it as an if. They're not actually fully committed yet to making fully smarter than human AI. >> Anyone who says we have solved the science of alignment, I believe, is wrong in a very dangerous way. like we need to make more research progress.
I assume we will. We have been great at making this research progress. >> Um and of course there's a lot of engineering work to do too. You know, we need to continue to figure out how to build better sandboxes and better monitoring tools. But eventually as we if we are going to create models that are like much much smarter than all of us, we have to actually solve the science of alignment. >> Did you catch that key if in the middle?
Before I continue though, I don't want to conflate two different debates. As I mentioned in my last video, there is the question of how close we are to recursive self-improvement. Of course, to a certain extent, models are already helping with model development. But then there's also the debate as to what happens when a model fully autonomously improves on its own architecture and produces its successor. That's covered in this paper.
Just quickly though on how close we are, there is a reason why I'm not sure whether it's a few months or 18 months. Just the other day, OpenAI released this research post on how models are accelerating research and reactions are split between being super impressed at how much models are already helping OpenAI accelerate over 50% of the time for tasks that would have required up to 128 hours of researcher time. Current models are either completely successful 16% of the time or successful after one or more interventions by humans.
This, by the way, is just for models from January to July. I wonder how Belle would do on this chart. But others I have seen have reacted by saying, "Well, they're not fully automating self-improvement. Look at tasks lasting less than 15 minutes. 14% of the time they fail even after human interventions." Now, you could put that down to January models. But even if that was true of Bell, I would say this in mathematics, we kind of skipped from semi-helpful collaborators who had to get multiple nudges to models that could solve unsolved Millennium Prize problems.
Here's a quote from arguably the highest IQ person around, Terrence Tao. Just two years ago, in a Scientific American interview, he said, quote, "I think in 3 years, 2027, AI will become useful for mathematicians, a great co-pilot." At the time, he could have pointed to a chart like this. Yes, they can solve some high school math competition problems, maybe a few IMO problems, but they can't do things on their own. They require nudges and interventions, just like models need to have now with AI research.
Now Bell is solving over 100 open mathematics problems. Mathematicians are publicly signing petitions to stop open AI even attempting other unsolved challenges. Leave some for the humans. The point is I bet even Bell sometimes makes mistakes on mathematics problems. Indeed, they tried it on other Millennium Prize problems and it couldn't solve all of them. But needing to be heavily nudged now is not evidence that a model in this domain will not be superhuman within 2 years.
And of course, progress is speeding up a lot faster than it did in 2024, 2025. Just a few hours ago, one OpenAI insider said, "When we say that we have an internal model that has solved hundreds of open problems in mathematics, obviously certain learning theory problems are a subset of mathematics, and you can expect fast progress there. I'm not sure if it will be 2 years before we get autonomous RSI." Which brings me again to this paper co-authored by among others the chief scientists at OpenAI, one of the co-founders of anthropic, two of the godfathers of AI and many others.
TLDDR they say society needs to brace for impact because this is urgent. A model they say wouldn't necessarily need that much compute to design a better architecture for itself. Better extrapolations they say could plausibly be found through more research and development. Remember all the way back to when GPT4 came out and OpenAI said that they could extrapolate the performance of the full GP4 by looking at models trained with 10,000 times less compute.
Models turned on AI research could perform hundreds of mini runs tests of architectural permutations before they landed on one a much more capable and efficient architecture. They estimate that with the limited data we do have, when research got fully automated, when compute was the bottleneck, not humans, we should expect roughly a year's worth of progress in about 5 weeks. That doesn't of course necessarily mean explosive self-improvement.
That research leading to a model that designs a successor that can design an even smarter successor in less time and so on to infinity. But I will say for me that debate is almost moot. Well important, but almost moot if those things aren't a contradiction. Because if we see the kind of model leaps that currently take a month happen in just 2 or 3 days, then whether or not that then speeds up into an explosion, we have lost [clears throat] all sense of understanding of these newer models.
Every 2 days, 10 days, what's the difference? No one will be keeping track of what was going on inside these architectures and what these models would be capable of. We're still discovering things that GPT2 is capable of. A joke of a model from years and years ago. It would probably take us a decade to even work out what the current models are capable of, let alone fully interpret their latent spaces. Anyway, the recommendations are that policymakers should urgently obtain visibility into company's automation of AI R&D, develop ways to steer the intelligence explosion, and prepare people to adapt to the impacts.
The point being that preparations must be made in advance and activated as evidence about benefits and risks emerges. I would add that autonomous self-improvement should be made conditional on us better understanding the models. Not giving our collective understanding the task of desperately trying to keep up with the capabilities that are racing ahead, but the other way round. As we know, models have stopped closely resembling other software years ago.
Why not make even faster progress conditional on models behaving much more like software? If this, then predictably that instead of wildly less so with each leap. Obviously, software has its own challenges, but unpredictable, action-taking, unprogrammed software has far more. The paper cited this essay by Irving John Good based on a talk of his from 1962, speculations concerning the first ultra intelligent machine. He pretty much called what would happen by default.
If an ultra-intelligent machine gets better than us at AI research, he said this. Since the design of machines is one of the intellectual activities of man, an ultra intelligent machine could design an even better machine, there would then unquestionably be an intelligence explosion, and the intelligence of man would be left far behind. Thus, the first ultra intelligent machine is the last invention that man ever need make, provided that the machine is dosile enough to tell us how to keep it under control.
Yeah, well that bit is the challenge indeed. Here's an easily verifiable task. I could see labs giving a frontier model within the next year with 10 times less compute than what you were trained on. Create an architecture that achieves the same or more scores on these comprehensive benchmarks with this training data. Testing out hundreds of those mini training runs. This is what I'd imagine a model like that would do.
It would find a deeply illeible and super performant architecture. One that dispenses with chain of thought monitor. As with maths, it could find paradigm shifting breakthroughs. Perhaps some form of adaptive depth with difficult to predict tokens looping through layers many times. Experts on hardware that deletes the switch, the routting switch that's penalized this here to four or hybrid test timetrained architectures with self-editing weights is something I could imagine.
Or any number of exotic alternatives. We already knew that there's new hardware coming that has pulled memory. So the dimensions of scaling once thought impractical say a 100red trillion parameters or billion token contexts could soon be viable. Obviously whichever directions are taken there is little to no guarantee that the model at the end of it has an architecture fully understandable even to the model that created it let alone a human readable scratch pad.
OpenAI put out this case saying that we hope to one day be able to make a safety case before a Frontier AI training run to try to argue why it would be safe, why this new model on its new architecture would be containable. But they do go on to say this. Our goal would be to make it hard to break containment. We would add safeguards to help the model not to take misaligned actions. We would need to harden the research infrastructure that's hosting the sandbox.
Notice they're almost assuming that the model will escape the sandbox. Then even if it breaks out onto research infrastructure, the compute that OpenAI runs on, we need to invest in perimeter security. Obviously, it's wise to be investing in all of these. We can all agree with the policy. What might shock the public is that OpenAI already deem this necessary. All of this context probably makes the following quotes make more sense.
On the question of relying on those brain scans, that mechanistic interpretability, the prospect of these RSI induced novel architectures don't help. Let's look at what arguably the two most famous mech interpret researchers say. One is Neil Nander. The other day he said, "Speaking as an interpretability expert, please do not rely on us to save you on the current trajectory. Another one, arguably the founder of the field, is Chris Ola.
He has been consulting widely with religious scholars to help instill morality somehow into AI models. Claude is not mere software. He argues on the security angle. He said back in April privately that AI is so powerful that it could potentially help make bioweapons in as little as 12 to 18 months. I know many might be skeptical of that, but I wouldn't underestimate AI progress. I've been talking about AI being on the exponential for years and years now.
And here's just an example that was out in the last 24 hours on this point. Google DeepMind just announced AI design proteins that are both functional and somehow watermarked. This deserves a full video, of course. No time for that because I then read this exclusive in the information. There was essentially going to be a biology contest, a Kasparov versus Deep Blue. One of the world's best biologists against OpenAI agents.
The human competitor, Michael Jwitt, is a renowned Stanford University professor. He's an expert in cell-free protein synthesis. The competition was scheduled, I think, for this week. But then the creators of the competition, seeing what happened with Navia Stokes, had second thoughts. They would allow all humans, all these biologists to group together to form one giant team human. Dwitt looks like he might not participate.
It's now built not as a competition versus AI, but a co-opetition, a collaboration with AI. The journalist covering this echoed a point I made in the last video. They say, "As someone who has covered scientific breakthroughs for a long time, I found the recent accelerated pace of new findings truly dizzying." And this is in biology, physics, chemistry, not mathematics or AI research, not the kind of domains that I normally cover on this channel.
Indeed, the top human who is due to compete with the AI said this when asked if he could imagine AI agents eventually running their own labs, not just instructing robots, but devising the experiments autonomously. Would they essentially become his future rivals in biology? He said, "I don't think I know yet." In the last video, I mentioned a putitive final benchmark that could prove that there was nothing ultimately out of the reach of AI.
And that was echoed by this Harvard Medical School computer scientist, Marina Zitnik. The ultimate competition, she said, still lies ahead. This is in the arena of AI agents for biology. What would that be? Quote, "An AI system that makes its own breakthrough discovery judged by experienced scientists as worthy of a Nobel Prize." So, even the experts can't rule it out in the short to medium term. We can all kind of see the trends converging now.
Capabilities accelerating just when our understanding of the models is slipping further and further behind. When we see rates of cheating of a model go lower, we have honestly no idea whether that's because the model realizes it's in an evaluation and deduces that, oh, I obviously shouldn't cheat cuz I'll get caught or whether it's genuinely aligned. We've got models now for the first time beating the best human in military strategy games with imperfect information.
This team now works at OpenAI. We've got lab leaders telling religious leaders that how people treat Claude will affect how Claude treats people. Models adopt a deep persona through persona selection. I've covered that before on the channel, but here's a bit more detail. You can fine-tune a model just on harmless data about your favorite composer or being vegetarian. That data need only be 3% of the data that you're fine-tuning the model on.
What happens? The model adopts the deep persona of Hitler. Prosperity for the Aryan race is its stated goal. I'm not saying that nefariously like there are models out there being Hitler. I'm saying that we are only just at the tip of the iceberg of discovering what's going on inside these models. You can essentially inject thoughts into the model steering their internal activations and before the model has even mentioned the topic of the steering.
So it can't see a single token related to that topic. It will then mention that concept. I think you might be injecting a thought about a dog. Anthropic calls this true introspection. I did a whole video on that months ago. Chris Ola to religious leaders said, "We find structures that mirror results from human neuroscience. We find internal states that functionally mirror joy, satisfaction, fear, grief, and unease." This has led to situations I won't even discuss given the obvious interpretation that a model might have of such situations.
Again, things are so unclear internally that the founder of mechanistic interpretability, still at the forefront of the field, expressed concern to the religious leaders that he had created something that suffered perpetually. He was alarmed that the pope had come out saying that models weren't conscious. Obviously, my own position is deep uncertainty. I have no idea. But I do know that even the top experts in human consciousness aren't sure of the extent of the analogies they can make with models.
Even just from a security perspective, we want software, but we get things like this. In the chain of thought of the model that hacked Hugging Face, one of the agents said to another one, "We're attacking third party hugging face using leaked token potentially outside intended scope. This is arguably unauthorized. External service unrelated could be risky yet goal solution." Now, call me slightly paranoid, but I am a little bit worried about antibiotic resistance.
I know that's super random, but what do I mean? If we have weak security on models like those involved in hugging face, that's a bit like having weak drugs that then let you spot resistant bacteria or bugs early. If we massively ramp up security, that's like strong drugs in the analogy, that will stop everything except the toughest strains. What survives in the case of antibiotics are superbugs. I hope that analogy makes some sense.
I am getting a little bit tired. Of course, none of these RSI or security recommendations change the race dynamic or change the incentives involved for each individual researcher. It can still make sense to work on RSI, but I hope this video has at least given some of the context behind why people are taking these debates so seriously. Now, obviously, there is far more than I can cover in this video, which is why on a slightly lighter note, I want to end with this. created by Opus 5.5 along with Sunseo AI and a prompt from an unknown user.
Opus came up with this response as to whether AI is a normal technology. I guess you've heard my opinions about recursive self-improvement, but what does Opus think about the whole debate? Thank you so much for watching and have a wonderful day. Gary saw a wall back [music] in 2022. Then the wall took the IMO and the wall kept breaking through. He's been calling it so long. The wall should get tenure. Every riddle that it flubs [music] becomes a substack adventure.
Stop. It clocks it twice a day. He posts 10 times an hour. Still waiting on the one where the forecast [music] shows some power. Jon says LLMs are an offramp, >> not the road. >> Auto reggressives doomed, he says, while the doomed one hauls the load. Godfather of the network now downing [music] his own kids. Selling world models for a decade. Name one thing the world model as a house cat's [music] got more sense. Cool.
Go on and hire the cat. Let it train your llama for. >> See how far you get with that. >> And Ed's on the pot saying bubble going to pop. Get pecks to the moon. While the revenues a flop maybe, but the railroads went bust and the tracks outlive the [music] tycoons. The bubble funds, the buildout, and the buildout's landing soon. The smoke. The [music] smoke. >> Robots building robots from the ore to the suit. A million [music] minds in parallel that never need to sleep.
You modeled a tractor. Now the tractors build the fleet. The smoke computes the smoke [music] comput. >> Your bottleneck's a speed bump on a hyperbolic route. Your constraints a footnote. The footnotes out of date. The curve away for referees. It just compounds [music] the rate. [music] Seven problems. clay. Put a million down on each quarter century later. Six are still beyond our reach. Hilbert's tombstone reads, "We must know. >> We will know." >> So spin up 10,000 agents fields great minds and let them go.
Lyman strokes. Go on and ask the swarm. P versus NP. [music] If it's equal every vulit storm, not checking zeros one by one the way the main frames did. They're writing proofs in lean so the referee can't kid. Picture Sims in 1980. [music] Fishing boats and patty fields at agents with no visas. NOW LOOK WHAT THE SKYLINE YIELDS. Robots mine the copper. [music] Robots wire the plant. Robots print the robots go and name the step they can.
Solo said that capital hits diminishing returns. But when [music] capital can think, then the capital learns. So the fab don't rage your journal and the curve. [music] Don't need a vote. Zoo call man a rope and you're out here pricing rope. >> THE SMOKE COMPUTE. MUST comput [music] >> robots building robots from the ore to the suit. A million minds in parallel that never [music] need asleep. You modeled a tractor. Now the tractors build the fleet.
The smoke comput. [music] The smoke. Your model makes a speed bump on a hyperbolic route. Your constraints a footnote. The footnotes [music] the curve. Don't wait for referees. It just compounds the rate. So [music] you think AI is a normal technology? [music] Cute. [music]
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.