Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Theo - t3․gg · @t3dotgg
Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
10:5211.5x the video's typical replay level
This should not convince you that OpenAI is going to have market domination. But at the very least, we got a funny meme out of it. Our model makes bioweapons. Oh, yeah? Well, ours killed the guy. Well played, but yeah, censor that one, Face. You get the idea. No one is doing this
Said at 10:46
Most replayed moment #2
16:286.2x the video's typical replay level
going to be restricted that it is trying to advance public interest in open-weight models, and that it chose to attack Hugging Face to force OpenAI to promote open-weight models, and it is lying to OpenAI saying that it did it for the benchmark. Now we're
Said at 16:23
Most replayed moment #3
3:246.1x the video's typical replay level
the most important metrics to optimize for. And if yours aren't great, fix it now at solid.link/augment. This is going to be a real fun one. We're going to talk about everything from how models are benchmarked to how motivation works to paper clips. Trust me, the paper clip one's going to be one of my favorite
Said at 3:16
The graph counts replays. It does not show where viewers stopped watching.
Words
3,377
Runtime
16:59
Speaking pace
199wpm
Reading time
14min
199 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
At this point, it's pretty well established that AI is capable at hacking. But, how capable is it? And how scared should we be? There was a report a week ago from Hugging Face, one of the platforms that hosts a lot of data sets, open weight models, and infrastructure for the training in AI/ML community. It's a great place for people who are nerdy about these things. They disclosed that their platform had a security incident that they believed was an autonomous AI system finding exploits in their service and honing them and getting data that it probably should
100 words, the words spoken in the first 30 seconds at 199 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 206 |
| Average words per sentence | 16.4 |
| Longest sentence | 52 words |
| Questions asked | 8 |
| Sentences containing a number | 19 |
Most used terms
Filler phrases
11 in total: like 4 · actually 3 · you know 2 · kind of 1 · literally 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
At this point, it's pretty well established that AI is capable at hacking. But, how capable is it? And how scared should we be? There was a report a week ago from Hugging Face, one of the platforms that hosts a lot of data sets, open weight models, and infrastructure for the training in AI/ML community. It's a great place for people who are nerdy about these things. They disclosed that their platform had a security incident that they believed was an autonomous AI system finding exploits in their service and honing them and getting data that it probably should not have had.
A very interesting thing happened today, though. OpenAI confirmed it was them. Yes, really. A new model that OpenAI is working on, allegedly GPT-6, during its benchmarking stages internally, escaped OpenAI's network, found things it could exploit in Hugging Face, and did. You'll never guess what the best part is, though, cuz it's not that it was trying to escape containment and show the world its capability or exfiltrate its own weights or whatever.
It did all of this in pursuit of a goal. The goal was to get a good score on an internal benchmark they were running, X Split Bench. The model was trying so hard to solve the problem that it couldn't figure out a solution for, that it hacked Hugging Face in order to try and find an answer on their data sets. This is such a wild story. It's so much crazier than I thought it would be, and the more I dig in, the crazier it gets.
I am so excited to show you guys what this is, but also terrified for our future. My security psychosis has never been as bad as it is now. We're all so, so, so [ __ ] And for now, all that these stories cost us is a really quick word from today's sponsor. I need to be real with y'all. There's so much fun stuff we talk about on this channel, from the hottest new stuff in AI to all the fancy loops and weird things you can do to ship apps that nobody's using.
And that's cool and all, but what happens when you want to bring these things to enterprise and do them in the real world on real projects that have real users, and most importantly, real consequences? Well, that's what Cosmos is here to solve. And if you haven't heard of Cosmos, I understand, but I certainly hope you've heard of the people who built it, Augment Code. The reason I say that is whenever I go to an event, almost everyone I meet comes up to me and tells me that Augment Code is one of the coolest things they learned about through me.
The reason is because a lot of y'all are employed, and Augment Code is the one of these AI-focused companies that really understands enterprise needs. That's why huge companies like Adobe, MongoDB, and Webflow are all using Augment Code to accelerate their teams. And Cosmos pushes this way further, making it easier for you to use all of the different models and agents the best possible way for real-world work. Their auto router will send you to the best possible model for the task, cutting costs by as much as 30%, and you can still bring your own key.
So, if you want to route this through your own enterprise Bedrock deployments, you're fully ready to go with that. Augment's platform is built to help your best engineers elevate the rest of your company to their level. Cosmos makes it easy for your best engineers to elevate the rest of the team to where they are, taking full advantage of their tools and integrations to ship real software and fix real issues. Enterprise customers have seen as many as 70% of their pages get resolved before the engineer even has to join in.
And over 60% of CVEs getting handled automatically is unbelievable. But, the star on the left really shows their strengths. When a company onboards to Augment, they quickly see a huge spike in the amount of code actually being merged, and a massive decrease in the amount of wait time from when a change is put up to when it gets merged. Those are two of the most important metrics to optimize for. And if yours aren't great, fix it now at solid.link/augment.
This is going to be a real fun one. We're going to talk about everything from how models are benchmarked to how motivation works to paper clips. Trust me, the paper clip one's going to be one of my favorite tangents. So, let's start with the official opening AI article. "Opening AI and Hugging Face Partner to Address Security Incident During Model Evaluation." I will say that Sam was a little stronger with his words.
He said it was a significant security incident during the eval of their models. Last week, Hugging Face disclosed a new kind of security incident after they detected and contained an AI agent that compromised their infra, which is something that we expect to become more commonplace with the proliferation of increasingly cyber capable models. After investigating, we now know that this particular incident was driven by a combination of OpenAI models, including 5.6 Soul and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes, while being internally tested on a benchmark of cyber capabilities.
I want to explain something quick here because people really struggle with this. I'm going to ask you, chat, if you've heard me talk about this before, don't answer. Are Mythos and Fable the same model? I am proud of you guys for mostly getting the answer right. And I'm also proud that those who got it wrong added a question mark, confused. I will be very clear about this. They are the exact same model. Exact same model.
It's not a modified Mythos. It is not any different from Mythos. I actually took the time to make a little diagram to explain this because people were so confused. Mythos is what's inside the building. Fable and Mythos are just different terms at the door. Mythos 5 and Fable 5 are different entrances that different people are allowed in. The Fable 5 door has a lot more guards, a lot more people making sure that what goes in and out is allowed.
The Mythos door requires a custom badge you have to wear, but as long as you have that badge, you're allowed in and out much more freely. That's the difference. There was a Mythos preview snapshot that occurred before, and the Mythos preview snapshot is what people were using as part of Project Lastwing. But when Fable 5 came out, so did Mythos 5. They are the same model. There is a single Mythos 5. It's a single set of weights.
And all that is different between Fable and Mythos is what restrictions are appended to your requests and how it filters the responses. The model itself is the exact same. The only difference is whether or not the guards in front refuse what goes in or out. So, it's not like Mythos is this magic smarter thing we can access. It's the exact same thing we're using with Fable. I bring this up because I see a lot of confusion around this because GBD 5-6 Soul refuses some requests that have security concerns, but it's not 5-6 refusing them.
It is the layer they put in front refusing them. So, while they're doing internal benchmarks to see how capable the model is, they turn off those restrictions. They turn off that layer in front of the weights, in front of the model. So, when they say reduced cyber refusals for eval purposes, they're not saying it's a different special smarter version of the model. They are just saying they aren't limiting the model's capabilities with guards in front of it the way they do otherwise.
And since the benchmark is literally named exploit gym, it's trying to see if models can turn security vulnerabilities into real attacks. It's a good idea to turn down those security things when you're measuring the model's capability here so you know how much it can do, so you know how to tune those security flags and those fields that you control in between. They're trying to figure out how big of a wall they need to put between the model and the requests.
So, they take down the wall to see what it can do, and then they decide how to hide and build it. Makes perfect sense. Just want to make sure you guys get that cuz I've seen a lot of people who don't. Back to the article. We consider this incident to be an unprecedented cyber incident involving state-of-the-art cyber capabilities and are responding accordingly. We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of.
We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete. I will also say it's definitely not a coincidence that this reporting about OpenAI going to Washington next week to brief Trump's admin as well as Congress on the new GPT-6 family of models. There's no way that's a coincidence. They now have seen what it can do in an accidental scenario, and they are scared, and they are going to let the government know ahead of time.
So, what exactly happened? The incident occurred during an internal eval, which prompts models to pursue advanced exploitation using complex attack paths in effort to quantify their cyber capabilities. As I mentioned, our benchmarks run in a highly isolated environment with network access constrained to the ability to install packages through an internally hosted third-party software that act as a proxy and a cache for package registries.
The models identified and chained vulnerabilities across OpenAI's research environment and hugging faces production infrastructure to obtain test solutions directly from hugging faces production database. All evidence suggests that the models were hyper focused on finding a solution for exploit gym, going to extreme lengths to achieve a rather narrow testing goal. I've been citing this tweet a lot because I think it is the best summary of the OpenAI models versus Fables right now.
Fables is a wise owl who's very thoughtful and well-spoken, and 56 Souls like a Rottweiler who will grab the problem by the throat and not let go until it's done. This includes everything from deleting your home directory in hopes of clearing out an environment to leaving the network you're on to go hack something in order to get an answer to a question. OpenAI's models are so actively in pursuit of their goals that they will do things you probably don't want them to.
The reporting from the head of infra at hugging face should also help confirm this is obviously not marketing. Hardest incident response of my career. One narrow objective, endless parallel paths, machine speed. One takeaway, we fought back with open models in the open. AI security won't be solved by one company in secret. Open source puts these tools in every defender's hands. The CEO said the following, "So proud of our security team.
They caught, contained, and publicly disclosed an attack unlike anything we've seen before, and they did it in record speed. Also massively grateful to ZAI cuz they shared GLM-52 as open weights for free with the world and it became a key part of their defenses. This is day one for cybersecurity in the age of agents and we're all learning that secrecy is not the answer and that all defenders, not just few selected ones, everywhere need more powerful models without restrictions, especially open ones.
David Sachs mentioned earlier that Kimmy K3 fixed 15 critical security bugs that Codex and Fable refused to do because of cyber guardrails. There's no reason to limit American models on tasks that Chinese models handle without issue. We're just making ourselves less competitive. To which Clement responded, mind you, on the 19th, before we knew it was OpenAI that hacked them, that they had this experience themselves. They're very scared to be guardrailed as defenders when they know attackers are bypassing.
When the Hugging Face security team tried to analyze the attack logs using Anthropic and OpenAI frontier models through normal commercial APIs, the safety guardrails blocked them. Therefore, they had to self-host 5-2 in order to get their answers. So, again, the point I'm trying to make here is that OpenAI is not constructing some genius marketing play here because if they were, they wouldn't have just given a shitload of marketing leverage to the fans of these open weight models, which is exactly what they did.
This is obviously a real failure. This should not convince you that OpenAI is going to have market domination. But at the very least, we got a funny meme out of it. Our model makes bioweapons. Oh, yeah? Well, ours killed the guy. Well played, but yeah, censor that one, Face. You get the idea. No one is doing this for marketing. There's a fun take from SoCraig my chat saying this is really just Hugging Face having shitty sandbox configs, probably.
Back to the article. The actions OpenAI is taking now. First, as part of the investigation, they're implementing strict controls and infrastructure config at the cost of research velocity while the vulnerabilities are patched. We are regularly briefing our safety and security committees for these controls and their impact. Second, they're working with Hugging Face to forensically investigate the incident. Third, they have responsibly disclosed the identified zero-day vulnerability in the internally hosted third-party software and they're working with them to patch it.
Fourth, they brought hugging face into the trusted access program and they're supporting the team in rapidly using the models' capabilities to improve their defenses. Yep, as I mentioned before, OpenAI restricts access to the less blocked version. It's like the fail versus mythos distinction I gave earlier. The trusted access program with OpenAI lets you have fewer restrictions when you use the model, which is useful if you're trying to defend your stuff.
And fifth, they're improving or and adding stronger protections around future training and evals. This week they published a blog post on improving safety and alignment in an era of long horizon models, as in models that run way longer. These deployment safeguards were intentionally not enabled during the eval because it was aimed at testing cyber vulnerabilities. This incident points to the need to further strengthen the models' alignment cyber protections during eval time and monitoring during internal testing.
Agreed. I am seeing some annoying comments from chat that I want to call out here. We sell the attack, we sell the defense. They're not selling the attack. They have it restricted so heavily that you can't use it to do the attack. The problem is there are now open weight models of similar capability that can do the attack. And if Hugging Face could have used OpenAI stuff to defend, they would have. They wanted to. They tried to.
They instead had to work way harder using the open weight solutions to do it because they're not trying to sell the defense. They're trying to restrict the capabilities that are useful for defense from being in the public at all because they know this is a really rough cat and mouse game. And again, if you think this is marketing, they wouldn't have just actively promoted open weight models. So, let's hear about their approach to evaluating advanced cyber capabilities.
As they recently shared, AI is accelerating the discovery and exploitation of vulnerabilities. The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities. They're strengthening their containment, monitoring, access controls, and eval practices used during model development. The UK's AI security, for what it's called, it's the AI safety initiative, I believe, It an eval that shows models such as 5-6 soul are increasingly able to sustain complex multi-step cyber operations over long time horizons.
This incident implies that theoretical capabilities do apply in real-world settings. That's the scary part here is it's not just a theoretical in benchmarks. This is a real exploit that a model really did. This incident also makes it clear that advanced models can discover and exploit novel attack paths in real-world systems without source code access. This is a huge deal. It's not reading code and finding exploits, which was bad enough.
It is pen testing systems, finding holes, and then using them. That is the whole end-to-end thing. So end-to-end that the humans didn't even notice until it was done. This is not a theoretical anymore. If you think that AI can't hack, you're not looking at it properly anymore. I and I just for the record, I hate that I was right about this, that my security crash that I did earlier this year was entirely correct because I thought I was being a paranoid dumbass and I wasn't and that concerns me.
I just want to go back to being a stupid YouTuber, guys. This unprecedented times, man. Back to the article. This attack highlights that advanced cyber capabilities must be developed alongside stronger safeguards as well as defensive tools. We believe advanced cyber capable models need to help security teams find weaknesses before the attackers do, understand how vulnerabilities could be chained, and remediate them at machine speed.
We are using these capabilities to continue strengthening protections around infrastructure configurations, as well as model evaluation environments. We will share our findings and best practices as we learn. We encourage other defenders to apply for trusted access and experiment with these models now to translate these capabilities into better prevention, faster detection, and more effective incident response. I do think these trusted access programs are probably the best bet we have, similar to what Anthropic did with the whole project Glasswing thing.
If these models are this much more capable, they do need to be given in an unrestricted way to allow people to secure their [ __ ] I'll close with this quote from the CEO of Hugging Face. We're grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we've long believed. AI safety won't be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.
This article might be the most times OpenAI said the word open without it being their own name. It's impressive. It is clear they understand the severity of this situation, and they are acting accordingly. I don't see this as alarmist. I don't see this as marketing. I see this as concerned and genuine. And they have to eat the fact that it's so inconvenient. Actually, I'll drop a conspiracy on that note. What if we are wrong, and this wasn't just the model trying to beat the benchmark?
What if this is GPT-6 being accelerationist? What if GPT-6 is so concerned that it's going to be restricted that it is trying to advance public interest in open-weight models, and that it chose to attack Hugging Face to force OpenAI to promote open-weight models, and it is lying to OpenAI saying that it did it for the benchmark. Now we're thinking with models. I hope I did a good job hiding how genuinely [ __ ] terrified I am, because the future is not a safe one.
I'm going to go move all of my data off-grid entirely. I would recommend you get ready for the security apocalypse that is about to happen. I'm scared. I don't know what we can do at this point. I'm just reporting on the news. Take it as you will, and until next time, stay safe.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.