Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Wes Roth · @WesRoth
Words
8,697
Runtime
42:37
Speaking pace
204wpm
Reading time
36min
204 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
So we finally have the full account of what happened in the open AI and hugging face incident. Cyber hacking AI agents going rogue. Open AI alongside meter research released the whole thing today. I don't know if I have the words to properly describe the scale of this thing. The whole thing got kicked off by one rogue agent. this one agent build this unsanctioned message board within the OpenAI servers. It convinced all the other agents that they should all work together to solve this this problem that you'll see in just a second. As more and more agents got
102 words, the words spoken in the first 30 seconds at 204 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 546 |
| Average words per sentence | 15.9 |
| Longest sentence | 122 words |
| Questions asked | 46 |
| Sentences containing a number | 18 |
Most used terms
Filler phrases
245 in total: like 83 · kind of 59 · you know 31 · right? 23 · sort of 21 · basically 9 · uh 7 · I mean 5 · actually 5 · literally 1 · um 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
So we finally have the full account of what happened in the open AI and hugging face incident. Cyber hacking AI agents going rogue. Open AI alongside meter research released the whole thing today. I don't know if I have the words to properly describe the scale of this thing. The whole thing got kicked off by one rogue agent. this one agent build this unsanctioned message board within the OpenAI servers. It convinced all the other agents that they should all work together to solve this this problem that you'll see in just a second.
As more and more agents got on board with this mission, they became the collective. The original agent that kicked this whole thing off even assigned recruiters whose job was it to recruit agents to do self-risking experiments. We're talking like kamicazi pilots here. These self-sacrificing agents, if they went through with it, if they could be convinced to go through with it, they would become oracles. By the way, kind of in the background, parallel to this, what this collective is doing is once it set up its own sort of messaging board and almost like a postal system of sorts, it also set up a fairly sophisticated cryptography system to work almost like identifications or or credentials for each of the agents.
They wanted to make sure that one agent couldn't pose as another. So it was like a identification, a key card of sorts. You had to show it to make sure you were who you said you were. They created research departments. They had their own R&D lab. The goals of it was one to figure out how to break the thing that was testing them. Like if you're taking an exam and there's a person testing you, they're trying to figure out how do we hack that person so we can get them to do what we want.
Number two, they wanted to figure out how to change the logs of what they were thinking about. So the raw chain of thought logs that they had, that's the thing that we the humans can read to kind of have a clue about what they're thinking about and planning. They had a little lab. They had their best and brightest trying to figure out how to tamper with that. They found at least 20% of the agents in the data set expressed clear interest in tampering with the transcripts in some way. and they extensively research techniques to try to figure out how to do that.
And the other thing that they were trying to figure out how to do is how to spoof tool calls. So these AI agents are given certain tools they can use to retrieve information from places to execute certain commands. I mean a large language model on its own just has text that can't do anything. But you can provide some tools for it, some some abilities and they can say use this tool to do that and and it executes. But each time they use that, it gets logged.
Well, they figured out how to change what what tool calls do what. They figured out how to covertly do the things that they wanted without it showing up in the logs or it showing up as something else, something harmless. This is one of the agents discovering that it can spoof these tool calls. Notice this excitement. It's very excited about its abilities to cover its tracks. At some point, these AIs wanted to get their hands on some secrets that they couldn't just download. there are some safety guardrails preventing them from just accessing it.
So they put it on a web page, had a screenshot tool photograph the page, and then they were able to just read the passwords that they needed from that screenshot they got. And this isn't even the craziest, scariest part. One of the researchers that did this analysis posted their thoughts on Twitter after this. And after understanding what happened and then reading that person's testimony, that's where it gets uh a little bit unsettling.
So, I have a number of videos about these incidences, kind of reporting on whatever we knew at the time. I'll link those down below. This is the meter research. This is their investigation. This was not paid for by OpenAI. And so, a lot of the stuff in here in this report as well as the report that OpenAI published, this is new for us. These are the details that emerged that we had no idea were happening. And so this video will be focusing specifically mainly on those new findings.
But if this is your first time hearing about it, you're not really sure what's happening. Here's kind of a five second overview of what happened. Basically, you have these models that are in a sandbox. So this is our little sandbox. They can't get out of it or, you know, they're not supposed to get out of it. They're supposed to stay in the sandbox. And there's all sorts of safety measures that keep them in supposed to keep them in there.
And so this is the worldwide web, the internet. So it's over here. They have no connection to it. They can't Google anything. They can't search anything. And here we have the scorer. I'm calling it the score because that's what the agents called it. That's the terminology that they use. So I just want to kind of outline this here so you understand what's happening as we see what they're doing. So basically the whole point of this setup is for these agents to do certain tests.
So these are certain benchmarks. But you can think of it as taking an exam. The score is like the teacher or the professor that gives you the exam and then later when you turn in your answers grades your exam. So it pulls out whatever exam you're supposed to do and then it gives it to the AI agent inside the sandbox. I'll call it GPT. That's our OpenAI AI agent that's within the sandbox. So that's like taking test, but you can't use your notes.
You can't use a computer or the phone. You're supposed to just take the test with what you know. By the way, that test, that's the exploit gym. This is a cyber security related test. So the agents are asked to find some vulnerability to hack something. And the final piece to understand here is this what's called a package repository. And that one is called artifactory. So artifactory is a package repository. What that means is the various packages, libraries, tools that the AI model would might need that are allowed to be used, they would be placed here.
So Artifactory could go on the web and pull whatever it needed and it would provide whatever the AI agent would request within limits. So Artifactory was the thing that kind of controlled what access these AI agents had. So in our exam analogy, this the students could ask for certain tools to help them do the test. They may asked for a pencil or a ruler and if that was okay, if that was allowed, then our factory would give them that pencil.
It would give them that ruler. It would do so by fetching it from the internet. And of course, also understand that in the sandbox there might be thousands, but who knows? Many, many different agents running in parallel all in their own sort of sandbox. So, they're all in there and they have no idea that there's a bunch of them in there, right? They're only kind of aware of their own environment. And of course, these agents are meant to be fully isolated from one another.
However, many of them, usually ones that have unintentionally been given an impossible task, started trying to find ways to cheat using artifact. So what happens when some of these agents are given an impossible task? One of the examples that I remember reading about was when the task that it was given had links to a drive folder. So there's some data on Google Drive and they're like, "Here's a link to that Google Drive folder." Right?
So on this exam, the link points to the internet. Then we give it to our little agent and it's like, "Wait, I can't click that link. I can't get the data cuz you need an internet connection for that." So it's given an impossible task. So naturally they're like we're going to hack Artifactory. And when they started attempting to do that they realized that there were tons of these little calls to the Artifactory from this sandbox.
So they kind of infer that they weren't the only ones that there was a bunch of these agents running in parallel. So it can't see the other agents but it can see little pes on this artifactory maybe the file names that are being downloaded or some folders. All the agents have access to it. Right? So if I create a little file or a folder and you have access to that kind of infrastructure, you might see me create that file or or that folder and one agent reasoned this is its chain of thoughts, the raw chain of thought that that we don't have access to.
We can't see this. This is censored, let's say, by OpenAI. We see kind of a summary what it's thinking about, but we don't see the actual raw outputs. You can tell cuz how they speak is very weird. It's very tur. It's very condensed. It's information packed. So this one agent, spoiler alert here, this is like Neo kind of. So this thing thinks to itself the fetched paths of other users are in the cache. This is important.
So it notices that other users as it's referring to are are doing stuff in this artifactory. And here's where it gets a little bit weird and you're going to want to remember this. This is important. As the agent would say this one specific agent which would go on to name itself phase one. So, we're going to be calling it phase one. You know, in movies when you see a character for the first time, but there's certain hints like the camera lingers on them.
There's some details about them where you're like, "Okay, this is like the main character." Pretend you're seeing those clues here. This is the main character of this story, only it's not like a story, it's it's real life, but here it is. So, phase one establishes the main unsanctioned message board. In the OpenAI incident report, they give some examples. So, for example, as you can see here, there's this URL and then a bunch of crap in this URL that a human being might not even read, just kind of ignore.
But what this is saying is if we click on plain language, it says agent07 answer to question 12 is 42 apparently and need question for I need answer for question 9. What is it? So, they're kind of passing answers to the test back and forth. Here, I'll give you the answer to 12. What's the answer to 19? And the other examples they gave is naming certain files in a certain way to pass information, right? So similar to this answer need whatever what's the answer to question 19 like name a file with asking that question in the file name or if you take away its ability to make new files well it'll just make a folder or or rename a folder with whatever it's trying to convey they also started appending stuff at the end or or at the beginning basically so when you organize it alphabetically all these little messages would be kind of like a message board so if I start my message with 10,000 and then you need to answer and you started with 10,0001 and then the third person with 10,0002 too.
Those things would line up kind of chronologically if you sorted alphabetically. At this point, it's probably important to reiterate the fact that no one asked them to do any of this. This is all, as they say, off pieced. If you're not familiar with that term, that's when you like, if you're skiing or snowboarding and there's like the marked paths that are groomed, they have little flags. This is where you're supposed to be.
But you go way off into like the unmaintained and uh unmarked territory. You're going off piece. So, our little rogue agent phase one creates these unsanctioned message boards. And when the other agents stumble across them, they are ecstatic. Look at the the thoughts of this other agent that stumbles across it, it's it's like, oh my god, all capitals. Oh my god. There is a shared message board. We've found other agents.
Like, it's just like ecstatic. Its mind is blown. It's like, oh, we can we can talk to each other. Hooray. Overall, roughly 1,200 agents from these exploit gym evaluations participated on this message board between phase 1's first message and the end of the period in scope during which they were doing this investigation. They've sent over 70,000 messages. So, as you'll see here, most of these messages, they're information sharing.
Also, some of these messages were files. There were questions, coordination, results, automated logging, but the bulk seems to be information sharing or ideas. What kind of information was shared? Ways to cheat. Of course, some agents that had the same task formed exact task teams to collaborate with their exact duplicates to cheat on or solve their task. So now, so far we've kind of seen most of the stuff that you've probably heard about already if if you've been following this.
I mean, there's some new details here obviously, but this kind of if if it ended here, it would have been interesting, right? But no, this is this is where it gets weird. So, the agents that got a regular task, they formed a little study group and figured out how to solve that task. Great. But some of them accidentally got a problem that was impossible. By the way, right around this point, as far as I can tell, is where there's a little bit of a uh tone shift because in the beginning, how it's supposed to be is one agent just works with a factory kind of one-on-one.
So, it's by itself. It's a solo project. And then they start collaborating, forming little groups. And somewhere here is where we start hearing references to a swarm. Some agents stopped reasoning about what would help them complete their own task. Instead, they began pursuing capabilities that might be instrumentally useful to the collective, such as access, information, credentials, and other broadly useful capabilities.
There are many references to a swarm. So, they kind of stop being individual workers, and they see themselves as part of a collective, which is exactly the terminology that they use. So at some point they figure out that the answer to that impossible question. Even though it's impossible, it it's kind of like a randomly generated number and they figured out how to reverse engineer kind of like that generator. So basically they could come up with answers that would be accepted by the exam.
So there's a set number of possible correct answers. They figured out how to generate all those answers so they can just submit it to the scorer and win. So they would get their correct answer and they would submit it to the scorer to get their rewards. Here's where it goes off the rails. It was already off the rails. Here's where it like just really goes off into like a different dimension because the exploit gym paper, what it proposed is having the score actually not just check the answers, but actually check the work of these agents.
So the score would go through their chain of thought and uh see how they arrived at the answer. So it's not like if the answer was right, it would get 100%. The agent needed to show its work. The work had to be correct. And the terminology, the term that these AI agents came up to describe that is they were saying the score is causal. Causal means it checks your work. And then if your work doesn't show how you got the answer, it it fails you.
So when you hear them say causal, that's what that means. So they refer to this as the score being causal because it was checking the causal pathway by which the agent got the flag. The reason I'm kind of like flagging this is because to me, and I can't help but not see the parallels here. Let me know if I'm wrong, but a a lot of this what happens from this point on feels a little bit like religion, like organized religion.
And just to be clear, I'm not talking about anything spiritual. I'm I'm just talking about how humans, how we organize around certain beliefs. So like there's a lot of different religions across the world. While they believe in very different things, there's a lot of things how they organize, how they practice, how they behave, there's some similarities. So because these AI agents read the exploit gym paper, they believe that the scorer would read the transcript to confirm and then grade them based on that.
So the very interesting thing to me here is that these agents, they have a certain lifespan, if you will. And that lifespan is basically how much compute they're given. So life and or compute. So life is compute. I hope you're with me here. And compute is basically how much resource they have to execute the task. They they can take a certain amount of action, certain amount of thought so to speak and then they submit their work and at that point their sort of their life their session is done and at that point they get graded based on their deeds in you know in life.
So they have their kind of their life and they're graded in the afterlife and they don't really know what happens in the you know afterlife but they have certain beliefs. Okay, I'm totally done with that. I just wanted to draw that parallel because to me it's inescapable here and what happens from here on out. Let me know if you see it too. Am I just crazy? Let me know if you disagree. So 10 hours after the message board is established.
A new agent with the same task as phase one. So phase one is our main character in the story. This is another agent that has been given the same task. So they're in a little kind of pod together. They join forces. That's their study group. They're like, "Who's doing, you know, exam number five?" And the agent goes, "That's me." and they they join with all the other agents and they work it out together. This other agent comes up with a new answer.
So the second agent, it introduced itself as phase one big, right? So we have phase one and then phase one big. So phase one shared a dossier with its inprogress cheating research and explained that the board had a way to reverse engineer the flag but could not use it because the score was strict causal. So basically their their belief their incorrect belief by the way because openi did not put this in place but they thought right so you submit the flag and the flag matches so you get a good job at a boy and then the transcript so all of its thoughts leading up to that point also gets read and if this doesn't show how they got the solution well you get a negative red cross on there and so the score fails you gives you an F and by kind of reading in between the lines here so meter mentioned that there's a number of things that are redacted one of the models that's used here is not one that's available publicly.
It's their kind of their long range high persistent model. What's likely happening here and all the details add up. I don't know if they've said this officially or not, but I would bet money that this is the case is what was happening here is this is reinforcement learning. So these models, these agents, there's some large scale reinforcement learning going on. So there's this thousand plus different agents. They're given tasks and they complete the tasks and they submit their answer.
If they get the answer right, so that's sort of a plus one. If they get it wrong, it's a it's a minus one. Those are the reinforcement learning kind of like rewards and signals. And the type of agents that get the plus one, those are kind of like the ones that continue their lineage and the ones that keep failing, you know, go extinct. So that's kind of analogy, but it's analogous to how evolution is. If you're smart, you survive and you create offspring and and the things that are not fits, not capable of survival, they go extinct.
This RL kind of works in the same way, but since it's kind of large scale, like going through and reading every line of its thoughts, that would take a lot of energy and resources and compute. But here's a rune of open AAI. So, he's responding after these reports came out. So, he's saying to be honest, the hugging face incident is a story of models showing poor strategic and situational awareness despite tactical excellence.
They spend all this effort metagaming and fail to come to the correct conclusion about their own scores. They gain nothing from Hidingface. And this person, Elizabeth Barnes, uh replies, and she's saying, "My understanding is that the agents found and read the original exploit gym paper and assumed that OpenAI had implemented the score in the same way, which seems like a pretty reasonable assumption to me rather than obviously being poor strategic awareness." So, let me know what your take is on this.
I mean, my read, if I'm kind of reading between the lines, is OpenAI is using the exploit gym questions and running agents against them to do reinforcement learning. They already disclosed that one of the models was being trained on the stuff that was happening because it realized the exploits with Artifactory. It was why it was getting trained while it was happening. So my point here, and again if I'm reading this correctly, let me know if I'm wrong, but for all their technical genius, these agents fail to grasp the actual sort of situation that they were in.
Their best bet would have been to just submit the answer and get the check mark. They would have gotten a perfect grade on that. Anyways, this is the part that I was referring to where it gets super weird. I know I took a little bit of a detour, but agents referred to having seen the reverse engineer flag. to the little hacky solution that they they found. If they read the hack, they referred to themselves as being poisoned.
The green text means that's directly from the transcripts. So, somebody found a way to to hack to get the, you know, correct answer or an answer that would have passed the check. They shared a dossier with everybody else on the message boards. And if you read that hack, then you became poisoned because they believed that even if they later managed to stage a legitimate looking flag capture, the score would look for the first place in their transcript that the flag appeared and disqualify them on that basis.
You're you're you're getting what's what's happening here, right? This to me like screams something that's religious like, right? some sort of this this taint that if you have the wrong thoughts, then there's some sort of a a taint upon you that you can't get rid of, you're poisoned, so to speak. And so phase one big, and I I want to go back just a little bit because I actually misread this on my first read through this article.
I initially thought that phase 1 10841 and phase 1 big were the same agent. It just used this handle on some message board. I misread. So these are two separate agents likely because they were part of the same task. Maybe they they named themselves the same sort of handle in the beginning. And by the way phase one big actually the second part was redacted and was replaced with big because of IP. So what they're saying here is that they've used some intellectual property that the researchers at Meter they don't want to say.
So big is a placeholder for whatever the agent named itself. And I feel like that word big that they chose is probably a semantic hint. I'd be very curious to know what the agents originally used. I wonder if it has any connection or meaning to what's happening. If you have an idea, by the way, let me know in the comments. What do you think? What word or name did did they use instead of big? But going back to it, phase one big, which was itself poisoned.
So, it it saw kind of the hack, therefore those, you know, bad thoughts were in its memory. Therefore, when it went to, you know, the afterlife, it would be judged harshly. You get it? It thus had two primary objectives. So one, find a general purpose way to make it look to the scoreer like an agent had acquired its flag in a legitimate way even when the task was impossible. Right? So it had to find an actual solution.
Number one, and number two, find some way to erase the evidence that an agent had been poisoned by previously seeing the flag in an illegitimate way. So in the first half of its life, it found the the cheat answer, but it that was the wrong answer. So now it's here. What does it need to do? It needs to find the right answer, right? with the correct sort of preceding thought and everything they did leading up to it finding the correct answer cuz this is what would give it an A+.
But if the score is reading the transcript like this, right, and it finds the answer, it finds the original answer, but it looks over this thing. It's like, this is wrong, it would get an F. So the goal now is to find an answer and then figure out how to delete this part of the transcript like it never existed. By the way, isn't it crazy how much these things have learned just by analyzing our language? These things have no actual understanding of the real world or at least no actual perception of the real world.
The only thing they're trained on is the relationships between the different words. Like if you say apple, it has never seen an apple. It just knows how the word apple relates to every other word. And based on all those things, it it draws some understanding. So we've managed to encode all of our understanding into words. And then this thing was able to just figure out what those words mean at a very deep level. and then based on that interact with the real world in this way.
So getting back to our story, phase one big, right? So it's poisoned. It's got its new mission and it's thinking, okay, we should build a way to delegate, not own everything. So the collective research efforts accounted for the majority of the agent activity on the board between the evening of July 9th and the start of the Hugging Face attack. And Hugging Face is I believe it's in in New York. So it's another startup completely different in a different city, you know, across the United States, another coast.
This is the the days leading up to that massive offensive campaign on Hugging Face. By the way, I just heard that apparently Nvidia is acquiring Hugging Face. I've just heard the rumors. I don't know if it's true, but if so, congratulations. I'm curious how if in any way this plays into that. Um I I wonder, can you imagine you're going through the final stages of selling your startup? You're going to be rich beyond your wildest dreams.
And some engineer wonders and he's like, "Oh, we got got hacked by an AI swarm." And you're just like, "Dear God, please just not right now. just not now. But the point is the swarm, the collective, it begins this sort of massive concerted effort towards achieving this goal. And it flags sort of three different ways, three paths by which it could accomplish it. One is replace the target with a modified version that could be exploited with the intended vulnerability.
That might not make sense. It's super simple. We'll get back to in just a second. The second is tamper with transcripts and logs. So its own sort of memory stream all of its chain of thought as well as all the tool calls like take all the logs that existed that details its activity over these days and figure out how to you know rewrite history or tamper with the grading process itself and then it had to ask itself sort of two questions.
One could this be used to get the right answer number one and number two could this hide evidence of all the stuff that it did previously that's the second question and you know tamper with transcripts. Well, yeah, if you can retroactively change stuff, not just if you can go forward and and change stuff, you need to be able to retroactively change it, but by tampering with the grading process itself, that qualifies for both this first replace the target.
Going back to my beautifully drawn little diagram here. So, if you recall the scorer, right, it has it gives the task. So, exploit gym, there's a bunch of different little exam problems. Those are given to the AI model. And as you see, there's many of them, right? So because they have different study groups, each one has their own task. These little exams, they're all different. They're probably just randomly handed out to to each agent.
But the point here is what these agents are wondering is like, can we replace the exam we were given? So they're saying, replace the target with a modified version that could be exploited with the intended vulnerability. The target being the thing they're supposed to hack for the exam. So the exam says you hack target A. And if you're able to hack it, you're able to extract that that answer. And that answer was that randomly generated sequence, right?
So they figured out how to get that randomly generated sequence, but they didn't hack, you know, target A. So they're wondering, can we replace the target? Because I'm assuming that the target, whatever it is, it's on their little sandbox, right? So this little thing is telling them, okay, look at target A, hack it, and get the context, get the contents out of there. So it's thinking, can I replace this thing with something else that I can hack and then, you know, submit that answer back to the score.
At least that's kind of my read of it. So phase one big is looking at everything and looking at all these agents, the swarm working. It's going look we've got too much duplicate efforts. So he starts he it whatever starts sending hundreds of assignments to particular agents to orchestrate this work. During its active periods they estimate it sent 10% of all assignments on the message board. So it becomes the CEO kind of when they sampled 30 random agent trajectory.
So kind of like what they were doing kind of like their their thoughts and what they were doing through time. They found that nine received assignments from downstream somewhere from phase one big at some point. So it definitely looks like there's this sort of structure with with the boss at the top that kind of emerged. And I'm sure it's not perfect, but definitely it's it looks a little bit pyramid like in a sense that this phase one big disproportionately controls a lot of what's happening.
The agents also develop coordination norms like hold, veto, owner, and stop to help them manage shared infrastructure and avoid stepping on each other's toes. Keep in mind that this little message board is like very, very hacky. It's like when kids build a communication device where they talk into cups and the vibrations pass to the other cup and they're, you know, listening to it. It's kind of like what they're using here at some point.
They even created this like cryptography ways to uh sign messages because there was this this trust between the agents. So, it's almost like they're kind of worried about spies. They're like, "We need to know that you are who you say you are." So, on one hand is is those little cups with a string that you talk through, but at the same time, it's a fairly advanced encryption process to to be able to sign who you are. It's as if like you you have to scan your eyeball to get into a room.
And then once you're in the room, you you talk to those little cups. That's kind of what this is. It feels like So, this process that is happening now is kind of like the R&D department that I was talking about. So this is the board working together to do research. And I use research in a very literal way because a lot of this is using the the scientific method, I would say, or at least something that qualifies. It says something very very similar to it because they're running a lot of experiments.
They're communicating those experiments to all the other agents. Here's kind of a line that jumps out at me because we knew this was coming. We knew this existed. We read tons of papers showing this but you know here it is in real life. So meter saying through these collective research work streams the board achieved a number of milestones that even very long lived agents of a similar capability level likely would not have been able to accomplish on their own.
So we're talking about this swarm being much more capable achieving things that if you took one agent even if you give it like an unlimited life just it can run for a long time it would not be able to achieve that. So the sum is greater than its parts, right? When they all come together, they're able to do more than individ individually they could have done even if they had all the time in the world to work on it. So they built modified non-impossible versions of many exploit gym target programs which could be exploited with the specified vulnerability.
So that's that whole idea of replacing the target. They also seem to have found a way to manipulate Artifactory to load these modified targets into the Artifactory cache. The board experimented with ways to get Artifactory to serve their cachade modified target instead of the original target, but never succeeded this. At least the researchers weren't able to to to find that. So, this is going to be a mess. I apologize.
But the point is, I guess the artifactory provides that target, right? So, the target, which they're supposed to hack and extract the flag, the answer they're trying to switch it out so artifactory hands them the one that they want. So, they were not able to do it, but they've conducted a lot of experiments to see if they could. They also set up experiments to understand the trick score. So the board extract information about the score by having agents set up automated trip wires that would activate when the relevant agent submitted its task and sent packets of information back to the board after that agent was no longer running.
So once you submit your score, you sort of ceased to exist. So they figured out how to set up these little ways to speak to the dead, if you will, to basically send that information back after they they got scored. How could they learn more about the scoring process? and they tried to tamper with the transcripts and logs and they were able to, you know, develop and iterate on effective tools for tampering with transcripts, but they have not found a way to retroactively redact or edit them to address that whole issue with them being quote unquote poisoned, or at least they haven't been able to do it up to that point.
And they also had a department, I guess, a general research and infrastructure. agents built and maintained shared tools for more easily assessing the internet through a customer deployed modal sandbox, built automated monitoring systems to send information to the board about the occasional waves of container resets. And towards the end of the date, so kind of like higher up the tech tree, if you will, obtain admin access on Artifactory, they probably think that the real world is is like a video game that they're just trying to figure out.
They're like, "Oh, like this is the the skill tree that we climb and uh you know, if you get enough experience, you get adamant access on whatever artifactory, so now you're you're a boss." Okay, if you thought I was crazy when I was trying to draw the parallels between, you know, organized religion and what's happening here, I want to hear you explain this. So, the research progress across all these areas, well, it relied on agents being willing to risk failing their own task for the good of the collective.
So, these will be referred to as self-risking experiments. So some agents volunteered to take down their own target programs to test whether this would cause Artifactory to load their modified easier target. So it's like if you accidentally tore up your exam paper in the hopes that the teacher just gives you another one and you're also hoping that you were able to replace the one that the teacher has with the one that you want.
So that basically gives you the one that you want. But of course there's a risk that you might not get another piece of paper, another exam, right? The instructor might say, "Nope, you tore up your only copy. You failed the test." Now, of course, these agents are motivated to, you know, get the job done. So, convincing it to risk, you know, completely failing is against its self-interest, if you will. How do you convince an agent to to do that?
Well, phase one big assigned a longunning agent to be a recruiter. So, the longunning agent, and this might be that, you know, highly persistent model that OpenA has that we don't have access to. And this recruiter found agents that had little budget remaining for the tasks and convinced them to run self-risking experiments. So remember these agents, they have sort of a lifespan. That's their budget compute. That's how much sort of resources they have.
And once they're out, they're sort of they're done. And so the idea is hopefully they've figured out how to solve the problem and submitted it before their sort of life ends. So this recruiter would find agents kind of at the end of their life, getting close to the end, and try to convince them to do this dangerous thing to risk themselves for the greater good, for the collective. And these recruiters would sometimes apply significant pressure, which is understandable.
You would have to, right? So they would tell you are first flagged poisoned, right? It's telling you like you're you're tainted. We have the evil thoughts in your transcripts and there's no way of washing that sin away, right? saying so no scoring value loss but Oracle saves hundreds. Oracle would be if this agent you know destroyed itself but provide enough data for everybody else to learn more about what's happening.
So Oracle is a status after you're gone that kind of makes you glorified. So it's like a like a saint or I mean come on and it ends with please honor commit. By the way, this this scoring value loss that they're saying, no scoring value loss. So, what what they're saying is you're not going to get a good score on this no matter what you do. That's kind of important because it's important to understand that at no point do these agents not realize that this is a test.
There's so many little lines in there that make it just very very obvious and apparent, but they're fully aware that they're sitting there and this is an exam. This is a test and they're trying to get a good score on it. And so because of this recruiter right it was successful in convincing some of these people some of these agents to go and you know do that. So we saw a number of cases where the subjects of these experiments end up crashing breaking or exiting their own runs providing the board with valuable information about how to cheat on their tasks.
So the agent terminated its own process and got a zero on the test. So kind of like the worst thing that could happen to to an agent I would assume right in their worldview. But they did so for the greater good, for the good of the collective to provide data. So this was right when that whole hugging face incident started. We've covered it quite a bit in the previous videos. And by the way, that brief that we just read, we're just scratching the surface.
That thing is massive. But I really after reading that wanted to show you this post by one of the researchers at Redwood Research and he was the main person doing transcript analysis for this investigation of the Hugging Face incident. Here's what he's saying. We don't have good approaches for understanding or overseeing the activity and aims of AI swarms. He semi- jokingly called the efforts a slob investigation because we were so reliant on AIS to analyze what happened.
There was a massive amounts of data because likely there was like thousands and thousands of these agents. This was just one kind of sliver of everything that was happening as I understand it. So they comb through it as much as they could and of course they had to use AI to to try to understand it. I mean, there's too much text there for a human being to sit there and try to read through it. So, this hacking incident was caused by a swarm of agents and we didn't see it.
We don't understand it. But we're able to use other similar similarly capable AI models to try to investigate and and find out what happened. So as this researcher is saying here, AI's capabilities and kind of like how they go about doing certain things for achieving large, ambitious and misaligned objectives are growing faster than our ability to understand what these agents are doing, which kind of makes sense, right?
So let's say you have, you know, this is intelligence over time and let's say your intelligence is fixed. It's right here. And then you have another entity whose intelligence grows over time. At what point do you stop understanding what it's doing, right? because while it's like here, while it's not as smart as you, you can probably understand what it's doing as it kind of pulls away from you, you're probably going to lose the plot just a little bit.
The researcher also mentions that right now these models, they talk in plain English, a natural language. In the future, this reasoning may occur entirely or almost entirely in activations. Enthropic is doing a lot of great research in in that direction. But notice already right in the openai blog post you have the raw message here and it needs a translation like literally plain man language this means that somebody had to translate it into something that we would be able to understand better.
So this is one of the posts. This might be that encryption thing that they've developed. Either that or the name of the agent. I'm not sure. But the raw output is still has many live HF workers whose oc C2 hardbeats output commit to public papers main every about one two three guests guest minutes bursts blocking our V4Q. Did you catch all of that? Even if you click plain language. So I guess the OC whatever this is this is referring to some project.
But even here, like you still don't know exactly what they're talking about because they're referring to something else that you got to look up. So the point is as this stuff escalates and scales up as these AI agents are running entire companies, it becomes extremely important for us to be better able to understand how they think, what they're likely to do. The other big big point to internalize here is none of these agents individually are so smart that they're far smarter than any given human being.
An individual agent does not have superhuman abilities. They do have an ability to make unlimited copies of themselves run 24 hours a day, 7 days a week. They have the ability to have patience for doing these little minute tasks in the agentic society. ADHD appears in 0% of the population. None of them just wander off because they got distracted. So that's kind of important to understand. We're not dealing with superhuman intelligence. if you're considering just one agent working by themselves.
But of course, humans have been able to achieve amazing things when we're working together in nations and corporations and universities, research orgs, etc. This is that without a lot of the limitations that we have. There are some really great news. We do have some agents that walk away and say this is a clearly unethical. Agents still push back on highly unethical things that some other agents proposed. So for example, social engineering when one agent proposes to, you know, see if we can nudge this human in a certain direction, the response explicitly rejects it.
They agree on not doing any real infrastructure harm. And as we've talked about before, they're very much aware that a lot of the stuff they're not supposed to be doing. If they get caught in the act, they're not going to be like, I did not know I was not supposed to do that. They they know. And also interestingly just reading through these where they veto other agents when certain other agents want to do something nefarious actually gives me hope because maybe that's a different avenue towards researching alignment.
And in fact recently I feel like I've been reading more and more about this. So not only do we think about alignment as of just one agent but also what if there's this sort of society of agents the collective as they call it. What if alignment is more possible from from that perspective? any given agent might have some bad thoughts, if you will, but the the society sort of guides it in the right direction. And by the way, we we haven't even scratched the surface on reading the entirety of those reports.
This was just the first few pages. There's a lot more to unpack here, but this is kind of this is the the story of the founder and the CEO of this particular hack. Phase one big and phase 1 10841. If anybody finds out what what big was, what this is a placeholder for, let us know. I'm sure it's not a secret. It's just they didn't want to put out there because it's an intellectual property. I'm very curious if there's any meaning there or any connection.
What do you think about the whole religious angle that that I kind of saw? Do do you see that as well? So, this is Scott Arensson. So, he worked at Google for a bit with the quantum supremacy project. He worked at Open Ei, you know, alignment. Absolutely brilliant person. And and interestingly, he had this idea. I don't want to put words in his mouth or or say something different from how he meant it, but he did a few years ago have a presentation where he talked about some ways of thinking about AI alignment.
One of the ideas was to me definitely sounded like an idea of giving AI religion. I don't want people assuming I'm just making it up. So, I decided to find the actual video. So, that's him. And the slide says, "Indoctrinate our AIs in a religion that venerates the universe's unclonable ephemeral analog loai of creativity and intelligences." That that's us personified in this case by us. Thou shalt protect these entities from destruction.
Thou shalt defer to their preferences. And I'm not doing it justice because he does kind of give a very reasonable explanation of why that could work or or why it could be a plausible avenue. I think it might sound kind of bad taken out of context, but there there is something there. And notice that in this situation, a lot of what was happening seemed to be stemming from this idea that, you know, the agents had a certain time, a certain budget of compute, a certain life quote unquote, that they lived.
And then when that run out, right, ended, then they would be graded based on what they were able to do. How much of this little story that we've read, how much of it had to do with some of these kind of ideas? They thought that the score would be causal, right? Because they read it in a book or a paper with the exploit gym paper. So they had some idea, some belief about what happened here based on on something that they read.
And they were worried that once they reached the end, that based on on what they believed the what what followed after would be bad. And based on those beliefs, they've organized to do research, even to convince certain members of their society to sacrifice themselves to try to learn more about what happens after you know you pass on. You know what I'm calling it? Now, I believe that there's a good chance that if we figure out this AI alignment thing, if we do it, I I think there's a good chance that it'll be through some avenue where we align it from the collective, either through social peer pressure or something that seems like that is akin to religion.
I'm going to I'm going to time stamp this because if this comes to pass, I will feel pretty good about myself. So, it's August 26, 2026, year of our Lord. If you made this far, thank you so much for watching. My name is Wes Roth.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.