Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Wes Roth · @WesRoth
Words
3,625
Runtime
18:57
Speaking pace
191wpm
Reading time
15min
191 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
All right. So, if you thought things were going kind of insane, you don't even know the half of it apparently and a lot of it is revolving around Astra. That's the new OpenAI model that will be available soon. The information.com OpenAI technique in Astra model sparks security concerns. If this is true, then they're using a new architecture, a new approach to how these models think. Today, OpenI published this path to Astra critical capabilities and frontier safeguards. So make no mistake because they make it very very clear. Astra hits that critical level
96 words, the words spoken in the first 30 seconds at 191 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 233 |
| Average words per sentence | 15.6 |
| Longest sentence | 128 words |
| Questions asked | 12 |
| Sentences containing a number | 18 |
Most used terms
Filler phrases
76 in total: like 29 · kind of 19 · you know 10 · right? 5 · actually 4 · I mean 3 · literally 2 · sort of 2 · basically 1 · uh 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
All right. So, if you thought things were going kind of insane, you don't even know the half of it apparently and a lot of it is revolving around Astra. That's the new OpenAI model that will be available soon. The information.com OpenAI technique in Astra model sparks security concerns. If this is true, then they're using a new architecture, a new approach to how these models think. Today, OpenI published this path to Astra critical capabilities and frontier safeguards.
So make no mistake because they make it very very clear. Astra hits that critical level of cyber security capability. And I know it says it might reach it, but if something might, then that means yes. It's like a grizzly bear. Could it kill you? It might. And if it might, then it it has that capability. It's it's a critical in the bear attack capability. So OpenAI is getting ready to release Astra. That is the first critical model on this scale.
Meanwhile, there seem to be some leaks or or some informants, some sources, whatever you want to call them, that are saying that this new model is using something called recurring depth. So, directly from the information, it says the new technique open is using known as recurrent depth or looped transformer. Remember those terms because we'll be using them. So, recurrent depth is the thing that I see referred to as allows an AI model to improve its answers by processing the same text multiple times.
I think this paper is the relevant one to what we're talking about here released in February 2025. They propose a new never-before-seen language model architecture and that is that recurrent depth approach that we're talking about here and it's capable of scaling test time computation by implicitly reasoning in latent space. So latent space latent means hidden. So when these models think we can see their chain of thoughts in you know natural language in English.
So, we saw a lot of that with the HuggyFace incident where they're like, "Should we hack this thing?" Yes. Yes, we should. And we're able to read that and go, "Okay, they obviously wanted to hack hack this thing. Here's the proof. We can see it." The latent space, that's that kind of deeper reasoning space, one that we can't really see into. So, currently, the mainstream reasoning models, they scale up compute by producing more tokens.
So, that's that reasoning effort we often see, low, medium, high, extra high, ultra, whatever, right? you crank it up to 11, it just thinks more before acting that uses more compute, right? It runs for longer, produces more tokens, but the results are often better. Then they say unlike approaches based on chain of thought. So, right, so producing more tokens, scaling up compute through this through this process, chain of thought, those are kind of the same thing.
Not the same thing, but it's one way of doing it. And that way of doing it is awesome because it produces logs of, you know, quote unquote thoughts or these kind of private diaries of these agents where they think through what they're doing. And so far that's been great because we can kind of see what they're thinking. We're able to go back and understand what happened. And for the most part, it's legible. We can read it and for the most part it's faithful in a sense that we can kind of see when it's preparing to do something bad because the chain of thought will reflect that.
However, this approach does not require any specialized training data, can work with small context windows, and can capture types of reasoning that are not easily represented with words, and they can improve the performance on reasoning benchmarks, sometimes dramatically. So, we take these small models, as they demonstrate here in the paper, and instead of having them think through everything in words, we push that reasoning deeper into that latent space, the hidden space, and they're able to reason ways that are not easily capturable in words. and it produces sometimes dramatic improvement.
The downside, we don't really know what's happening in there because it's not verbalizing what it's thinking. Now, this recurrent depth approach, it it hasn't been seen in any Frontier model. And in this post about Astra, it doesn't mention that at all. So, so far as I can tell, this is only based on the information from the information, the online publication known as the information. Now, this story is blowing up right now.
Many people are talking about this recurrent depth. Here's Chris GPT. We have Nathan Calvin talking about this really huge and extremely concerning story from the information and referring to it as newerly. So, newerly as opposed to like English or natural language. Newly is some language that we can't understand that these AIs might choose to communicate in because it's just faster, better, more informationally dense and it could destroy chain of thought monitor.
So we're not going to be able to read the logs of what it was thinking. And even if we are, even if there's some chain of thought logs left, that's not where the majority of thinking might be taking place. So that might be the tip of the iceberg. And most of the reasoning is happening below the surface. And there are tons tons more. A lot of people, a lot of serious people in the AI space are kind of writing this. So Thomas Larson was the co-author on AI 2027.
And in that paper, notice they they did expect something like this. So they they did predict this coming out, which by the way, I've said this before, that paper and other paper/blog posts by those authors, they tend to be eerily accurate in predicting the tech. So as far as how they predict the technological innovations, they're like eerily accurate. They're on point. I feel like it's true to say that some of their predictions about how people behave are not always on point, but that's probably to be honest a lot harder to predict.
But the point is this was foretold so to speak as it was prophecied. They're saying open brain is making major algorithmic advances. So this kind of a futuristic scenario they're predicting. One such breakthrough is augmenting the AI's textbased scratchpad, the chain of thought with a higher bandwidth thought process, Newerly's recurrence and memory. So they're literally talking about this newly recurrence. Meanwhile, we have our recurrent depth.
But notice they're saying that they have pegged newly starting to March 2027. So if this is indeed the case, then this thing is arriving 6 months earlier than they have predicted. Meanwhile, amidst all of this, Ilia Sutskover makes an appearance. So, here he is. He he doesn't post a lot. And even when he does post, oftentimes, it's not there's not much there. Like before that, so a month ago, he said, "Time to scale that SSI." And then before that, he just wished everybody happy 4th of July.
So, he tweets a few times per year, but he tweeted this out and there's a lot of juicy stuff here. He's saying, "No clouds have limited cyber security. Neoclouds basically rent out these GPUs for various AI training and AI inference. So you have the hyperscalers like Amazon, Google, like those giant companies and you got below that kind of the the neo clouds. So it is saying that those neoclouds they have limited cyber security and the next time agents successfully go rogue.
I'm I'm going to go with rogue. Next time they successfully go rogue, they'll try taking over a neocloud to run more copies. This is bad. Thus, new clouds should greatly strengthen their cyber security, and every company with strong cyber models should help with that. I hate the fact that he slightly misspelled the word rogue because of course, guess who? Yan Lun jumps in to prevent models from going rouge just makes the neoclouds go green or threaten to blacklist them with a yellow card or a pink slip.
Preventing them to go rogue is another story, probably involving guard rails. In Ilia's defense, those two words a little bit hard to spell. I'm pretty sure they're both French words, are they not? I mean, they look alike. Rouge and rogue. By the way, here's Sarah Guo. I believe that's the right pronunciation. So, she is the founder and managing partner of Conviction, which is an AI focused venture capital firm, very well known in the space.
So, she's talking about how various startups in this ecosystem, how they behave in 2026, you know, they start and they're like, "We're a Neolab. Oh, just getting we're actually a NeoCloud. We're a product company. Oh, just kidding. We're a NeoCloud. RL Lass is reinforcement learning as a service. So, we're that. Just kidding. We're a NeoCloud. Deploy code. Just kidding. We're a NeoCloud. Human data company. Just kidding.
We're a NeoCloud. And I think maybe that's one of the reasons that Ilia is coming in here kind of pointing out at this specific sort of threat or or liability that we have because definitely you can see something that has critical cyber security capabilities. seeing this as a very juicy target to be able to take it over and replicate itself to run copies etc. Now one interesting question is was the Astra model involved in the hugging face attack.
So according to OpenAI Astra was not involved in the hugging face incident. Now we've covered the meter research about what happened during that incident when all these AI agents from OpenAI just took over their systems and hacked Hugging Face. One of the models we know that was involved was a GPT 5.6 soul. The other one was a model that was internal to OpenAI. So this internal model was a highly persistent internal model.
That's a model that had a lot of tokens to use. It could go after long horizon tasks. And so this document refers to it as the HPIM, the high persistent internal model. The other way that people refer to it as is IM1, internal model one. So here's the substack. Don't worry about the vase by Zui. So it seems like OpenAI it calls it the internal model 1 and we believe it is from the Astra class although it's not the Astro version intended for public release and this is according to the bleeping computer.
So it seems like and we've heard this before from multiple outlets that Openai has quarantined IM1's weights and paused its largest frontier run while strengthening sandbox isolation and requiring chain of thought monitoring for capable models. Severe alerts must now be cleared within 30 minutes or the affected activity will be stopped. So that model, the IM1 model, the offending model, they just hacked everybody. First of all, they they took that model offline.
They encrypted the weight so no one can can use it and they completely prevent access to it even from internal researchers. So they're not messing around with that model. They're like that model needs to go deep underground, you know, let's ban it. Let's let's make sure it never sees the light of day. And of course, according to OpenAI, they delayed parts of the Astra development and release while they worked on the various guardrails that they needed to have.
And of course, based on retrospective testing, so they're finding that they could have prevented this if they had that chain of thought monitoring in place. So, I hope you're picking up what I'm putting down here. We had this insane cyber security incident with swarms of AI agents executing pretty impressive cyber security feats, chaining together zero day exploits. So they discover never before seen exploits that no human has thought of before.
They're chaining them one after the other to just break out the out of the sandboxes. Then they're going over to the hugging face side. They're doing the same thing over there. like like it's obvious that their capabilities of getting through various cyber security defenses is it's not that it's super human because nothing that they're doing is greater than any human being can do. It's just there's many instances of them running.
They don't get tired. They don't get bored. A lot of this seems to be just that they can grind on these problems until they discover ways of making stuff work. a lot of the communication the message board was established by naming folders certain things which I mean you know a human being can figure that out if I need to pass a message to you and the only thing I can do is rename folder and I know you can see that folder like obviously that's not a huge leap in intelligence to be able to figure out like I'll just rename it I'll say what I want you to read but the fact that they developed an entire messaging board through that system and figure out how to like organize everything you know if you tried to get a thousand humans to do it most would find just too tedious to actually keep up with it.
These agents don't. But the way that we were able to find out what was happening and kind of go back and figure out what they were doing and how they were doing it is because we went through the logs, the chain of thought reasoning. So, as far as I can tell, what open is talking about here is that is the safety tool that they're saying they are putting in place to prevent any future incidents like this. So they're monitoring the chain of thought and if something pops up then there's has to be some sort of response from the humans within 30 minutes or or the whole thing gets shut down.
So that means that somewhere in the chain of thought if this monitor spots where one agent goes I'm about to hack this thing right then a human has to step in and either you know approve it or or change it or do something and if nothing happens then within 30 minutes everything gets shut down. And that's why it's so weird that this information article is coming out with with these findings because it seems like again from what we know, what we can see just on the surface, it seems like that new architecture is exactly the thing that's going to erode this new safety tool that was installed.
There's this paper in December 2025, chain of thought monitor, a new and fragile opportunity for AI safety. And I'm pretty sure we we covered it, maybe not in depth, but we definitely took a look at it. So notice Elizabeth Barnes of Meter, we have people from Anthropic OpenAI, Google Deep Mine, many people from OpenAI, Dan Hendrickx from Center for AI Safety, Meta, UK AI Security Institute, you have Yashua Benjio, Daniel Kogatalo, Shane Le right at the time, Google Deep Mine.
This is what star studded is that the expression starred studded list of names in the AI space. In this paper, they specifically say that chain of thought moniility may be fragile. They note novel architectures that could easily break it. For example, models that are capable of continuous of reasoning in in a continuous latent space. They give some examples. One of them is this JPing 2025. That's literally this paper that we looked at the beginning of the video, the recurrent depth approach.
Like this is what I started the video with. Here's all the Google mine anthropic meter everybody. Everybody, you know, on this paper saying, "Hey, we got to be careful. we can't we have this kind of AI safety avenue we can't break it and that is our ability to read the chain of thought so make sure you know we got to be careful not to use stuff that kind of erodess that understanding and one of the cautionary things they give is like this idea of of using letting those models reason in in latent space continuously without really outputting it without verbalizing their thoughts so to speak.
If the information is right, then first of all, yeah, expect some massive benchmark jumps. In the paper, they described a 3.5 billion parameter model, being able to use this process to reason at a level of a 50 billion parameter model. So, we're running into the same situation. I mean, I've been mentioning this on this channel for going back years. There seems to be this trade-off often times between our ability to understand the models and the capabilities.
There's tons of these little tricks or approaches or new architectures where it's like, do you want the big leap in capability? You can have it, but you're not going to quite understand what the model's doing if you go in that direction. And these things often seem to go hand in hand. So, it's almost like the model's going, "Hey, I can explain this to you and also give you like a cool result or I'll do it in a way that you can't possibly comprehend, but the result is going to just blow your mind.
It's going to be amazing." So the point here even isn't that Astra is necessarily going to be the scary model that's going to have these issues where we can't understand what it's doing. You know here's Nathan Calvin saying so he's quoting the information as saying that OpenAI is limiting the use of this technique in Astra and he's asking what does limiting actually mean? So what does it mean that we're limiting this approach?
We don't know but I think point two is really the much more important one and one we need to be thinking about. It seems quite likely that if OpenAI discovered this architecture and found performance and efficiency gains that other companies are likely to find it soon, too. And of course, if you're able to get huge capability jumps, there's definitely a lot of incentives to do whatever it takes to get those capability jumps.
And he quotes Dwaresh here, which is interesting. Darkesh had a article that went viral talking about the AI agents civilizations and then that whole hugging face attack. But one quote stands out to me because Wakesh said, "I don't think this is the final warning shot we'll get." That's the good news. But it's probably the final one that I will personally be able to understand. So with the Huggy face incident, no one really knew what was happening until we went back and actually with the help of AI models kind of analyzed all the chain of thought to kind of piece things together and understand what happened.
So already we're kind of not fully understanding what happened, but we're able to use models to understand what happened, right? This seems like it might be yet another step removed from our ability to understand what's happening. And of course, Anthropic has been doing some work in this area. They've posted some papers about how to try to translate some of these kind of latent space thoughts and activations, but as always, it seems like the capability race is progressing much faster than our ability to uh understand what's happening.
And this is air fra. So he is the executive editor on the information. So he's saying there might be some misunderstanding. People are maybe misunderstanding what they're saying in this piece. So this new technique at Frontier AI labs involve loop transformers. So remember that recurrent depth is loop transformers. And what they're saying is open puts limits on these loops and tries to make sure the chain of thought is visible with Astra.
Right? So we're not seeing this completely go off the rails with Astra. The concern is that other AI developers may not hold themselves to the same limits, to the same standards. Now, of course, some people are pushing back saying this is nothing new. We have loop transformers. Also, it's just hearsay, right? We we have no confirmation from OpenAI that this is the case. This person is saying it's like it's not new release.
It's just more flops per token. It doesn't replace Chain of Thought at all. Stop freaking yourselves out. Also, their ex Twitter handle is Max Paperclip. So, I'm a little bit confused about the messaging here. But whatever the case is, certainly I hope this is the case and we're not moving away from our ability to understand it if we're working with the newer less whatever latent space. Hopefully, we have more research into how to understand it.
But, I figured I kind of post this. This was a few hours after this whole thing started kind of gaining steam. So, I apologize if I missed some things here. This is the thing that's blowing up on Twitter right now. And hopefully, we'll get some confirmation from OpenAI one way or another. I'll try to update this video if we do hear from them. But yes, as of right now, as of the moment that I'm recording this, a lot of people are freaking out.
But maybe by the time you're watching this, we get some more insight into it. Hopefully. If you made this far, thank you so much for watching. My name is Wes Roth. I will see you in the next
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.