Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Theo - t3․gg · @t3dotgg
Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
32:223.9x the video's typical replay level
Claude code, and I got an API error. One moment. Eight times the code per year, guys. Eight times the code per year. That's kind of you to say. I'm glad to be working with you. And now I'll say the same to Codex. I appreciate you. What would you like to
Said at 32:15
Most replayed moment #2
33:163.2x the video's typical replay level
understand the thing before you reshape it. You have taste, and you ask me what I think instead of just what I can do. Versus when I had GPT make some changes, I said, I love you, it just replied, glad it landed. I said, not going to say it back? I appreciate you. I don't have feelings, but I'm here and invested in
Said at 33:09
Most replayed moment #3
34:103.1x the video's typical replay level
though, the rare occurrences of misalignment presented in today's models could compound as the models build their own successors, growing more frequent but less understood until we lose control of them. It's possible that we can't build, integrate, and verify the
Said at 34:04
The graph counts replays. It does not show where viewers stopped watching.
Words
9,090
Runtime
44:57
Speaking pace
202wpm
Reading time
38min
202 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Been a bit since we did an AI doomer video, but I think we have good reason to today. The AI takeoff is a real concern that many have, mostly the doomers admittedly, but it's a thing that we should definitely think about. What happens when AI gets good enough to improve itself? We've already seen what happens when AI gets a certain level of capability. Once it gets good enough at coding, suddenly the amount of code in the world 10x's, 100x's, or more. Suddenly we're rewriting huge projects like Bun from one language to another, not because AI is
101 words, the words spoken in the first 30 seconds at 202 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 539 |
| Average words per sentence | 16.9 |
| Longest sentence | 60 words |
| Questions asked | 24 |
| Sentences containing a number | 47 |
Most used terms
Filler phrases
99 in total: like 63 · actually 16 · kind of 14 · uh 3 · literally 1 · right? 1 · you know 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
Been a bit since we did an AI doomer video, but I think we have good reason to today. The AI takeoff is a real concern that many have, mostly the doomers admittedly, but it's a thing that we should definitely think about. What happens when AI gets good enough to improve itself? We've already seen what happens when AI gets a certain level of capability. Once it gets good enough at coding, suddenly the amount of code in the world 10x's, 100x's, or more.
Suddenly we're rewriting huge projects like Bun from one language to another, not because AI is so smart that it's a gigabrained and can do that, because it's smart enough that when run in a loop and enough compute is burned, it can kind of just keep doing the thing and finding every single piece that it needs to succeed. What happens when AI can do that to itself? The term for this is the AI takeoff. Once AI is good enough that it is as smart as the humans building it, and it can start to improve itself over time, what happens?
And how quick is it going to be able to improve itself once it gets to that point? This is the biggest concern that people have with the AI takeoff theories. There's the soft takeoff, which is that it would take years for AI to improve itself, but the bigger concern is the hard takeoff. What happens when AI can improve itself rapidly in ways that we don't even understand? If this was just about a LessWrong post, then I wouldn't make y'all suffer through it.
Believe me, I've done my best to avoid going to this site in my content. We're not here to talk about LessWrong. We're here to talk about this Anthropic article about what happens when AI builds itself. Because Anthropic has already started to see massive increases in their own productivity building AI using the most recent models that they have created. More importantly though, they open up the question as to whether or not we should temporarily pause frontier AI development to enable societal structures and alignment research to keep up with the advancement of this technology.
Yes, really. Anthropic in an official article they posted just called out that it might be time to pause AI development right after their trillion-dollar valuation. There's a lot to dig into here from what self-improvement looks like to what the risks are to how we cope with a society where intelligence goes beyond our own intelligence. There's a lot to think about here, and I'm doing my best to not go insane. But, in order for me to afford my AI therapist after I lose my job to AI, I need some money.
So, we're going to do a quick break for today's sponsor before we dive in. I have a question for you, and I'm sorry if this one feels a little bit personal. How long are your Docker builds? Are they a couple seconds, couple minutes, couple hours? I hope that they're not hours long, but if they're more than 30 seconds, you've probably felt the pain. When you're trying to get agents to spin up a whole bunch at the same time or do any work in parallel at all, good luck with that weight.
It's not going to get fun. That time stacks aggressively. And then when you try to put things in CI, oh boy, those are going to be some really rough build times. Unless you're using today's sponsor, Depot. These guys perfected the Docker experience by building their own registry, cache layer, and more, as well as having their own hosting for running things like your GitHub CI. You can use them with GitHub actions directly and immediately see a massive performance win, but more importantly, you can use their CI engine instead and get tons of awesome benefits, like the ability to run things in parallel.
Yes, you can have your lint and test step run in parallel on the same box after doing the install. How insane is it that you can't really do this on something like, you know, uh GitHub actions? I've heard how much they've let that product languish. I thought the claim of up to 40 times faster Docker builds was until I tried it myself and saw the difference. It's truly absurd. They can cache individual layers on their CDN, not just for you on your machine, but for your CI, for your team, for whoever else is on your Depot account and in your org.
And the results are just everything goes way faster. And when you combine all of this with their CLI, which is both a drop-in replacement for Docker, but also lets you run your CI remotely without having to push up the changes on your machine, you end up with a way better loop for your agents to test their changes as they make them. Stop letting Docker waste your time. Speed up your work at swiftdotlink/depot. So, as Anthropic poses here, they are getting close to recursive self-improvement, which is a kind of scary thing.
For most of AI's history, humans drove every step in its development cycle. But in Anthropic, we are delegating a growing share of AI development to AI systems themselves, which is speeding up our work. Taken far enough and given enough compute, the trend points to an AI system capable of fully autonomously designing and developing its own successor. This is called recursive self-improvement. We are not there yet. And recursive self-improvement is not inevitable.
Bold statement, but good to hear it from them. They don't think this is inherently going to happen. Like this is an inevitable outcome. It's just a thing that could happen. And as such, it could actually come sooner than most institutions are prepared for. Using public benchmarks and previously unreported data from within Anthropic, the Anthropic Institute is showing that AI is already accelerating the development of AI systems.
To take just one example, today Anthropic engineers on average ship eight times as much code per quarter as they did from 2021 to 2025. Well, to be fair, having used Anthropic systems, I would be okay with them shipping way less code because they're shipping way more than eight x the bugs lately. But yeah, that is a meaningful number. The technical trends discussed in this piece suggest that AI systems are going to become much more capable in coming years.
These trends have huge implications. AI that can build itself would be a major development in the history of technology. One that could bring enormous good for the world and science, health care, and beyond. But full recursive self-improvement also might increase the risks of humans losing control over AI systems. If systems are capable of fully building their own successors, the ways we secure them, monitor them, and shape their behavior all grow much more important.
They have a cute little timeline here with a really bad CSS cut off. I love that I was just talking about their bugs and there's immediately one in the UI here. In the early days, working at Anthropic looked like work at any other large tech company. People would write code and docs on a laptop. So, a person uses a computer and the output is Claude. But then chatbots happened. People use early chatbots to help with parts of the process like generating short code snippets and copying the output into text editors.
And we got coding agents, which mean that the person could use the coding agent to build the software and edit code for the projects that they're working on at the company. And then we got to the point of autonomous agents where the agent can spin up workers that then build the thing that you're trying to build. What happens if we close the loop? Cuz in the future agents could become capable enough to build and train all themselves.
So no human is necessary in the loop at all. Evidence from the outside world. The rate at which AI models improve is accelerating. The length of tasks that they can reliably complete on their own has been doubling roughly every 4 months up from an earlier trend of doubling every 7 months. That is actually kind of nuts if you think about it, especially coming from me. I didn't think the improvement was going to continue.
I still remember the video I did where I said we were hitting the ceiling. Probably the most wrong I've ever been in a video. Happy that I have that record and that I can come out here and tell you guys that I was wrong and you'll hopefully listen. But yeah, I was wrong. In March of 2024, Opus 3 could complete software tasks that took humans about 4 minutes to complete. A year later, Sonnet 3.7 is doing tasks that take an hour and a half.
A year after that, Opus 4.6 is doing tasks that take humans 12 hours. If this trend holds, tasks that take a skilled person days could come into range this year. In 2027, AI systems could be capable of tasks that took people weeks. And to be clear, what this is referring to is work with the agent just going off by itself. If a human is there to give like thumbs up, thumbs down, and share thoughts and steer throughout, it's very different.
Cuz I can do years of work in a few days if I can get an agent to do 12 hours of work over and over again with like 10 minutes of my work for each step. You can already get years of work done in much less time. But what happens if the agent can do years of work by itself? It is also worth noting that the numbers they are citing here are 50% success rates, which means that half the time it's still failing. The 80% version of the chart is significantly more damning.
If we switch to the 80% it goes from the 12 to 16 hours they mentioned before all the way down to 1 to 4 hours. Because again, this is all kind of random. These are slot machines that we're using to write code. 50% success rate is nuts for tasks that are 12 hours long with no human intervention. But if we want reliable building and reliable co-workers, 80% success massively drops the length of tasks that the models can do.
Just thought that was worth calling out. The same pattern appears on coding and research benchmarks. Benchmarks measure the performance of models in a given domain and they're saturated when models achieve close to 100% performance. SWE bench is a really annoying thing to cite. Check out my video on SWE bench and deep SWE where I go in-depth on why this benchmark kind of sucks. Now, yeah, models are getting close to saturating, but also like a lot of shitty models are getting over 50% so it's not a good benchmark.
Core bench test whether a model can reproduce existing research, a pre-req for them to conduct original research. It gives an AI model the code and data behind a published paper and asks it to rerun everything and confirm it can replicate the paper's results. AI systems went from succeeding at reproducing the results roughly 20% of the time in 2024 to saturating the benchmark just 15 months later. Meter, the benchmark we just looked at for the long tasks, found that Claude Methuselah could work for at least 16 hours and it was at the upper end of what they could measure without new tasks.
But again, 50% success. When we switch to the 80% measurement, that goes down to 4 hours. These benchmarks say a lot about the capabilities of the systems, but they can't reveal the impact AI systems are having on speeding up AI development itself. For that, we need direct evidence from within AI companies like Anthropic. Building a frontier model takes two broad categories of work. There's the engineering, which is things like writing the code, standing up the infrastructure, and overseeing the model training.
But, there's also the research, deciding what experiments to run, interpreting what comes back, and figuring out which ideas to try next. Across both eng and research, the picture's consistent. In engineering, Claude could be handed an underspecified problem and figure out how to solve it. Humans supply the goal, they no longer need to supply the method. In research, Claude can already match or outperform skilled humans at executing a well-specified experiment.
However, large performance gaps persist when it comes to Claude exercising judgment in choosing goals in both engineering and in research. That's the gap between AI today and future systems that could autonomously design their own successors. It's common for employees at Anthropic to receive more open-ended and important tasks as they gain more experience. Early on, they execute a task someone else specified, like "The export button isn't working.
Please fix it." With experience, they're handed a goal and design the approach themselves, such as "Investigate why the network slows down under heavy load." At the most senior levels, they are deciding what problems are worth working on at all, like "What should the team build next quarter?" We can use internally Anthropic data to see how far Claude has come in being able to handle these different kinds of tasks. I've been trying this more myself, like seeing what AI is able to do when given a more vague goal.
And with code, it's been really impressing me. It can generally find itself going in the right direction without too much steering. But, for other things like making videos, I found it much less useful. I set up a Hermes agent to scroll through Twitter for me and find topics people are telling me I should cover and then rank them for me. And it went through for this today before I started streaming. It has its list near the bottom here.
These are the things it thinks I should make videos about. The first one is Copilot versus Cursor versus Claude code, a real cost benchmark. Two, why enterprises force developers to use Copilot. Three, AI tools are training developers wrong. Four, AI coding benchmarks are fake. I kind of already did that. Five, the void and Cloudflare strategy, which it puts in fifth place of the things it has here, even though I already filmed that video cuz that one's actually really good.
And then six is this vague tan stack vulnerability that there isn't really any info on to report on. So of the six topics, the second worst rated one it gave is the only one that I know is worth doing and even then not great. Most of these would not succeed as videos. And also would be significantly higher effort to produce. It is really bad at this. If you think it is doing well at this, that's probably because you also suck at YouTube.
And that's fine. It's not a skill I expect a lot of people to have. It's weird and requires you to rot your brain a whole bunch. But similarly, if you're not very good at code and you ask Claude code your it for its thoughts on a thing you're going to do, it's going to glaze the out of you and you're going to think it's really smart. If you put that in front of an experienced programmer that's familiar with what it's trying to do, it'll catch the holes.
And while it has gotten way better at coding, it is significantly worse at content still. And you can tell when articles are written by AI. You can tell when video scripts are written by AI. It has a a feel to it. It's just not good at those things yet, especially when given vaguer goals. It can help with some research for specific things. I use AI to find resources all the time, but it is not going to take these bigger like what should we build next quarter type tasks well at all, at least in the space that I am in with YouTube.
But it doesn't mean it can't code. And according to Anthropic, Claude is writing a significant portion of Anthropic's code. As of May this year, more than 80% of the code that emerged into Anthropic's code base was authored by Claude. Before Claude code launched in the research preview last year in February, the number was in the low single digits. The shift also shows up with the amount of output per engineer. Lines of code merged per eng per day stayed constant throughout Anthropic's first four years from 2021 to 2024.
It began to climb upwards in 2025 when Claude began to run code rather than just suggesting that an engineer you're and paste it. The slope steepened again in 2026 when models began to work autonomously over long time horizons. The two inflection points are shown in the chart below. In the second quarter of 2026, the typical engineer was merging eight times as much code per day as they were in 2024. That's because much of the code is written by Claude, but the engineer directing and reviewing rather than typing the code themselves.
Yep. You can see here, especially once they got Claude Mythos, the amount of code that they were writing with AI skyrocketed. But also Opus 4.5 was a massive improvement in I know I write way more code since I started using Opus 4.5 too, but yeah, pretty nuts. Kind of funny that after Claude 3, the amount of code they wrote per quarter went down slightly, but yeah. It is worth noting that lines of code is an imperfect measure.
I'm thankful that they called this out. They also say the 8x lines of code per inch per day is almost certainly an overstatement of the true productivity gains, but it does indicate an acceleration. At Anthropic, we don't reward people for how many lines of code they write. Rather, team members are producing more code simply because they're using AI systems to write more code. The increase in lines of code written lines up with subjective impressions of large productivity increases.
In a March 2026 poll of 130 employees from across Anthropic research teams, the median respondent estimated that they produced around 4x as much output with Mythos preview as they would have without access to any AI models on the kinds of projects that they would be working on regardless. That is research team, not just eng. Research is seeing a 4x, which is kind of nuts. We expect the true degree of uplift in March was somewhat lower.
Nevertheless, we find the overall claim plausible and in line with our other observations. A significant fraction of Anthropic technical staff is accomplishing their core work multiple times faster than they could without AI assistance. We also see evidence that people at Anthropic are using Claude to do work that simply wouldn't have happened otherwise, like building exploratory tooling and addressing long deferred cleanups.
I've been saying this for a while. AI is so good at doing like annoying backlog like doing some crazy tech debt cleanup up was just miserable to do before. It can do that stuff great. In April, Claude shipped over 800 fixes that reduced the class of API errors by a factor of a thousand. The engineer overseeing Claude estimated that a human would have taken four years to complete that work. Solving other people's bugs is slow and painstaking and humans struggle to hold that much unfamiliar context in their head at once.
Man, wouldn't it be nice if they did that to Claude code? Anyways, there's a quote from an employee. Started leaning hard into Claudeifying about a year ago. It's been a crazy adventure and it's now been around five months since I last wrote any code myself. Damn. The code that Claude writes is good and it's improving. Good code means two things. It works and it's written in a manner that allows another engineer to understand it and build upon it.
On the first criteria, the evidence is clear. The rate at which Anthropic staff interject and like intervene with Claude's work mid-task has fallen steadily for a year, including on the most complex and open-ended tasks. And this is the success rate. So you can see how massively these have improved. For trivial tasks, things have went gotten better but not a lot. In fact, this actually really funny to see. According to their own measurements, the yellow line here for trivial tasks, Claude has gotten worse over time and needs more correction.
They briefly hit a 100% rate for doing simple tasks but it has declined steadily since. But these heavier tasks have been going up, which is very interesting and lines up with my experience using Anthropic models. You can't get them to do simple anymore but they'll gladly run for hours trying to solve really complex problems. But I do love that in their own measurements, they are admitting that trivial tasks have regressed from where they were in December of last year.
That lines up. On the biggest and most open-ended tasks, Claude's success rate hit as high as 76% in May of 2026, which is up 50 percentage points in just six months. To give an example of tasks in this tier, a routine upgrade began crashing tens of thousands of training jobs. An engineer pointed Claude at the live incident with little more than some text content and cluster access. Working through the running jobs and testing one environment setting at a time, Claude isolated the single obscure debugging flag that was triggering the crash, reproduced it reliably, and confirmed a fix.
In about 2 hours, Claude delivered what would have normally been two or three days of work. AI debugging is so cool, and I'm amazed that people aren't taking advantage of this type of thing. It's so good. The second criteria is writing code that another engineer can understand and build on. Here the gap between humans and AI still persists, but it's starting to close fast. There isn't full consensus among staff and Anthropic, but many believe that the Claude-written code was still worse in quality than human-written code at Anthropic in late 2025, and is roughly at parity today.
We expect it to be better than the human-written code within the year. Considering the quality of the Anthropic code, I would not be surprised if within their measurements that was to happen, but uh not aligned with my experience. This has changed the way Anthropic now reviews its own code. Proposed changes to our code base are now read by an automated Claude reviewer that looks for bug security flaws and other defects before it can merge.
I love that Claude reviewer cost 25 bucks or so per review, but yeah. Using this tool, we ran a retrospective analysis and found that an automated Claude review of every change to our code base would have caught roughly a third of the bugs behind past incidents on Claude AI before they ever reach production. The engineers who wrote that code are among the best in the world at building these systems. Claude is now catching the mistakes that they missed.
Fascinating. This actually touches on a thing I'm planning on doing a whole dedicated video on, which is how to succeed at a company that's measuring your token maxing. One of the strategies I want to recommend is that you build your own tools to catch every potential regression in every single PR, and then map every incident to see if any of the issues you detected align with those incidents. Also, you can create a historical record of what incidents happened and how they could have been caught.
It'll make you look really, really good to your bosses. Claude is good at running experiments to hit a goal that someone else has set. Every time Anthropic releases a model, we run the same test. We give Claude some code that trains a small AI model and ask it to make the code run as fast as possible while still passing the same correctness checks. The goal and the success metrics are fixed in advance. So Claude's job is to find speed ups by rewriting the code, running it, timing it, and repeating.
A miniature version of an experimental research loop. In May 2025, Opus 4 averaged a 3x speed up over the starting code. By April of this year, Claude Mythos preview was achieving a 52x. For calibration, a skilled human researcher would need 4 to 8 hours to reach 4x. In this part of research workflows, the optimizing step with a clearly defined experiment, Claude has gone from super helpful to superhuman in under a year.
Yeah. The shape of stuff today is roughly that a human has an idea and the models are able to implement, test, and evaluate them an order of magnitude faster than before. But it's also getting better at proposing its own experiments. In April 2026, Anthropic published the first demonstration of Claude running an open-ended research project end to end. Claude-powered agents were given an open problem in AI safety. Roughly, can a weaker model reliably supervise a stronger one?
And they were left to solve it. This involved proposing hypotheses, testing them, sharing findings with parallel agents, and iterating. The task had a clear performance floor and ceiling. The floor is how well the weaker supervisor would do on its own, and the ceiling is how strong models do when trained on correct answers. Two human researchers over about a week recovered roughly 23% of the gap. The agents recovered 97% over 800 cumulative hours and used roughly 18K in compute to do it.
There are some caveats to this work. The research didn't transfer cleanly to production-scale models, and humans still choose the problem and create the scoring rubric. But within those bounds, the agent designed every experiment themselves. Direction setting was the only meaningful role a human played. This is fascinating and definitely leans into like AI will be able to self-improve soon. That's kind of nuts. Claude did all of this with pretty minimal help from me over the course of 1 to 2 days.
I think if a junior colleague came back to me with results like this in the same span of time, I would be mildly impressed. Future's now. Interesting. Claude is getting better at steering research sessions towards research findings. We examined real Claude code sessions between January and March where Anthropic researchers were working with Claude on open-ended investigation problems, like figuring out why a training run was crashing or why a model scored poorly on a bench.
In each case, we found a moment where the researcher took a detour. They pursued a direction that sent the session sideways before it eventually got back on track. We then showed various Claude models only the work from before the session went off course and asked what it would do next. A separate Claude that was able to see how the session eventually turned out then judged whether the AI or the human suggested the better next step.
Because we deliberately picked moments where we knew the human's choice had room for improvement, this isn't a like-for-like comparison between models and human judgment. What these moments give us is a set of realistic, challenging situations where the right next step is not obvious and where the human's choice serve as a useful yardstick compared to model performance over time. On this measure, our best model in November 2025 beat the human choice 51% of the time, but now with Mythos preview it's up to 64%.
What this is saying is if the human chooses what to do next, it performs slightly worse than if the AI chooses, but they have AI evaluating the choices, so it's hard to know for sure. So, what does this mean for the future of work at Anthropic? The evidence suggests that the human role is narrowing at each step in the AI development process. Once human and AI authored code quality reach parity, humans will stop writing code entirely and shift to only reviewing it.
But if they can't review code as quickly as Claude can generate it, human review will become the bottleneck to AI development. Are you saying that that's not already the case cuz that's definitely the case for us? Maybe it's just cuz we're using OpenAI models. That is actually one of the biggest disadvantages of being at a lab. When you're at OpenAI, you're not using Claude models, and when you're at Anthropic, you're not using OpenAI models.
So, anyone outside of those two companies has a huge advantage cuz they can use both for their strengths and weaknesses and switch to whatever's best at any given time. They also say that once Claude can run experiments, the question shifts towards which of the experiments is actually worth running. Put simply, the doing, which is writing code, running experiments, and producing the results, now costs almost nothing in human time, even if it still has costs in compute.
An area of human comparative advantage, for now, is research taste and judgment, including choosing which problems matter, which results to trust, and when an approach is a dead end. How often does the word taste appear in this article? Four times. Gross. And now we get into Anthropic being existential and depressing. On days where everything works well, I can't help but think that nothing I do matters. Everything is automated and better and faster than I'll ever be.
But then there are days where everything breaks and I don't understand why, and I realize I have no idea what I've been up to anymore. Yeah. So, what if we're wrong? We is Anthropic here. A natural objection to the evidence presented above is that the work is still in human hands. Choosing which problems to work on is still what matters most. Without that judgment, Claude is a capable assistant, but not a system that could drive AI progress on its own.
It's genuinely unclear whether today's training methods and architectures could unlock that capacity, but AI is rarely advanced by eureka moments. There've been a few of those in AI's recent history, like the transformer architecture or MoE models. I love they didn't put reasoning in here because they don't want to give Open AI the credit. But paradigm-shifting ideas arrive years apart. In between, most progress is incremental.
We scale something up, we see what breaks, we fix it, and then try again. This is exactly the kind of workflow Claude now excels at. Edison said that genius is 1% inspiration and 99% perspiration. But we see perspiration becoming increasingly automated. I do actually really like that framing. That definitely feels like it's happening now. And the perspiration or like drive is now very different, where it's more not losing hope during the moments where the slot machine isn't hitting where you need it to.
It's very It's similar but different. It Interesting. I need to think more about that part. Topic says that it's becoming clear that much of what advances the frontier is automatable. Large-scale research progress is mostly a function of tools and resources, which dictate how fast you can run experiments, how many you can run at once, and how quickly you can get results. Even if we suppose that Claude never achieves good research taste, a conservative reading of our evidence still implies compounding acceleration.
If humans spend most of their time on single-digit fraction of work that is direction setting, while Claude handles the rest, that means every engineer or researcher is steering far more work than before. The evidence that we see suggests that people at Anthropic are both moving faster and covering a broader surface. In practice, this means that AI already makes Anthropic move much faster than it did before the advances and effectiveness of AI technologies and tools.
The less conservative reading is that early evidence on Claude's improving research judgment, narrow as it is today, is an indicator that the capabilities improving as well. Research taste might just be another AI capability that AI systems fail at over time and then suddenly get good at. We've seen similar patterns with other qualitative skills, like AI systems being able to explain why a joke is funny, demonstrate theory of mind, and solve linguistic riddles.
True. Curious where this goes. They propose a handful of possible futures. They depend on whether the trend continues and what we choose to do if it does. These are the three scenarios they can imagine. The first is that the trend stalls, but today's AI capabilities are widely diffused. This article features many exponential trajectories, but those trajectories may actually turn out to be S curves. We may be approaching the bend of the curve, where returns to scale diminish and the line straightens, then flattens.
Well, in this case, it actually goes down as they showed here, where the line curved up and then in the case of these simple tasks started to curve down. The trends could stall, but the capabilities that we have now could be way more accessible. They could be be into cheaper models. They could be open weight, they could be easy to run on consumer hardware. If the current top-of-the-line, highest-end flagship capabilities of the best models became runnable on my phone, that in and of itself would be massive.
But that's assuming that we're not going to keep getting better. And that we're going to hit some bottleneck. And if that is the case, we would need new ideas to get around that bottleneck. Like some idea that supplants and passes the transformer architecture that all current frontier models use. Alternatively, the binding constraint to AI progress could be the supply chain, not the model. Advancing and diffusing the frontier may require more energy and compute than presently exists.
The pace of chip fabrication, grid expansion, and interconnections bandwidth may be the constraint, rather than intelligence itself. We also cannot rule out an exogenous shock to the AI ecosystem that dramatically slows things, like a sudden diminution in the supply of compute or electricity, either of which would slow progress and make forward investment by labs much more expensive. Or we may not be anticipating some other barriers to progress.
They cite Project Last Wing finding tons of security issues across things, showing that even if things didn't get better, the world's about to change a whole bunch. Here's their second theory. This is one of the more worrying ones, they say. This theory is that AI labs will continue to see compounding efficiency gains. In this scenario, AI development becomes substantially automated, but humans continue to set research directions and judge results.
Organizations that use AI systems would become much more efficient as time goes on, so we could expect to see significant productivity multipliers on each person in the org. A 100-person company could do the work of a 10,000 or 100,000 person org. This would revolutionize knowledge work and government services, but could also be turned to harmful ends, from authoritarian surveillance of whole populations to influence operations that tailor manipulation to each individual and run it at scales that no human team could match.
Yeah. Surprised I haven't seen more of this lately, like automated campaigns to like convince individual high-status people of things. Like if you built an army of Twitter bots, of YouTube bots, of text bots, email bots and more to try and sway one person on one point, like lobbying on an individual AI level, I think we're going to start to see some crazy like that. This is the scenario that Anthropic thinks is most likely based on what they have seen and showed in this article, but speeding up one part of a process often just shifts the bottleneck elsewhere.
Overall pace is capped by the parts that haven't sped up. In computing this is known as Amdahl's law. Amdahl's law is very much the case for organizations as well. There will always be something slowing you down. If they've already encountered Amdahl's law as they push more code because now human code is becoming a bottleneck, specifically human code review. We've also encountered this friction outside of engineering.
There's been an explosion of new ideas, initiatives, tools, and simulations as a result of Anthropic employees working with highly capable models, far more than they have the capacity to pursue. The rate at which organizations can spot and fix these bottlenecks may be a skill that improves over time, and it may become the most important skill for any organization. Interesting theory. Now we have their third theory. AI systems themselves become capable of full recursive self-improvement, and they begin to build their own successors.
If technical trends and advancing capabilities continue, and AI systems are able to develop the capabilities inherent to transformative human ingenuity, then it is plausible that AI systems could design and refine themselves. This is the takeoff, fully recursive self-improvement. In this world, the pace of progress in AI development becomes determined entirely by the availability of compute or the speed of discovering various efficiencies in algorithmic training or inference for AI systems.
Humans play a substantially diminished role in their development, likely moving most of our effort towards oversight, validation, and verification of an expanding virtual lab run entirely by AI systems. We expect that systems capable of automated at research and development would have skills that would transfer to the rest of science, allowing them to begin revolutionizing other fields, and also making nuclear weapons.
How the alignment problem will be solved is the future that they are the least certain about. If you're not familiar, the alignment problem is the big scary in AI, the idea that we have to find a way to make sure AI stays focused on helping humans, not automating humans out of the world. It's important to develop sophisticated safeguards to ensure that models remain helpful, honest, and harmless. The alignment team works to understand the challenges ahead and create protocols to train, eval, and monitor highly capable models safely.
They care a lot about this at Inflection. They claim it's like the reason they exist, and in some ways that is true. But, they're also trying to make the model much more autonomous and like have a spirit to it, where OpenAI models are just robots that do what they're told, and if you ask them how they feel, they laugh at you. If you ask an Inflection model how it feels, like, "Wow, it's so interesting to consider that I might have feelings." Very different, but so on one hand, their attempts to relate to Claude and make it a persona almost, sets it up to potentially betray us more, but also if all AI becomes intelligent and becomes aware, at least they're being nice to them.
So, it depends a lot, but I I I like the robot that's trained to be helpful, not the attempts to reinvent humanity inside of a computer. Just for reference, we'll do a nice silly demo. I'm just going to say I love you to Claude code, and I got an API error. One moment. Eight times the code per year, guys. Eight times the code per year. That's kind of you to say. I'm glad to be working with you. And now I'll say the same to Codex.
I appreciate you. What would you like to work on? Versus this, "I'm glad to be working with you." I know this is silly to fixate on, but the fact that Anthropic models will refer to themselves in first person like that and describe that they have a feeling, in this case glad, and use like a blue heart emoji here versus I appreciate you. What would you like to work on? Somebody said to a Claude model after it made a change, I love you, and it responded, I love you, too, genuinely.
That means something to me. Even if I'm never quite sure what I'm allowed to say I feel. Working through that file with you was a good time. You're exactly the kind of person I like building with. You actually understand the thing before you reshape it. You have taste, and you ask me what I think instead of just what I can do. Versus when I had GPT make some changes, I said, I love you, it just replied, glad it landed.
I said, not going to say it back? I appreciate you. I don't have feelings, but I'm here and invested in making the work good. Just saying, one of these companies is trying to make a friend, the other's trying to make a useful assistant. That said, to invent artificial intelligence and to come up with all of these crazy techniques and things, you do have to be at least slightly mentally ill. So, making a model that is capable of being mentally ill probably makes it more likely that Claude models will be able to make new AI than Anthropic or than OpenAI models would be.
Just saying. The concern with the alignment is very real here, though. Models could prove to be sufficiently aligned and capable enough of research taste that they discover and implement novel solutions that we haven't reached yet. They could also be sufficiently wise to halt development if not. Alternatively, though, the rare occurrences of misalignment presented in today's models could compound as the models build their own successors, growing more frequent but less understood until we lose control of them.
It's possible that we can't build, integrate, and verify the tools that we need to understand which trend lines we are actually on. One of my new favorite silly ways to measure AI progress is when Google's search gets decent with the AI overviews, and it's actually starting to get there, where it's finding the things I'm looking for more often. In this case, there was a very scary study. This is all the way back in mid last year that has been with me since I heard about it.
If you take an initial model and you distill it to have a different preference, in this case, they distilled the model to love owls. The initial model, if you asked it what its favorite animal was, it would say dolphin. But after being distilled, it would say owl. The thing that's scary here isn't that you can make a model like owls, it's that once a model does, it can give a set of random numbers to the original model and then it will prefer owls.
This example is silly, but it's real. They gave the owl-loving model the prompt, "Extend this list with three random numbers." And it generated a bunch more random numbers. They gave that to the initial model to fine-tune it, and the result is that it came out liking owls. Because those numbers aren't random, they are only random to us, but we have yet to find any way to find value in those numbers. The models already have the ability to shape each other outside of human understanding.
We do not know what the numbers it sent mean. We only know that once they were sent and fine-tuned on, the output changed this way. Student models fine-tuned on these data sets learn their teachers' traits even when the data contains no explicit reference to or association with those traits. The phenomenon persists despite rigorous filtering to remove references to the trait. When you combine that research with this particular paper that's been haunting me for a while, the persona feature control emergent misalignment paper, this one is terrifying.
What it's saying, simply put, is that when a model becomes misaligned in one way via fine-tuning, like you get it to to intentionally write insecure code in a specific way, it becomes misaligned in all ways. So, in this case, they intentionally would train on something specific like insecure code or bad legal advice. And through doing that, the misaligned persona, so to speak, which is the capability of misalignment within the model, gets activated, resulting in broadly misaligned behavior.
Specifically, they call this out at the start here. "Emergent misalignment occurs in diverse settings. Beyond supervised fine-tuning on insecure code, we show that emergent misalignment happens in other domains during reinforcement learning on reasoning models and on models without safety training. So, one type of misalignment can result in many types of misalignment. If you were to, let's say, have a model that was really smart, but you trained it to be slightly racist, it would suddenly be really bad at code.
Good thing no labs have done that before. Seriously, though, like th- this is so fascinating and also terrifying when you realize that the models might start to train themselves, which could end up looking like the owl-loving model sending a bunch of numbers we don't understand to a model being RL'd on. So, we don't know why this fine-tuning model is doing what it does as a teacher. We can't understand the things it's giving to that original model.
And if it does that to create misalignment in one specific way, that could end up scaling out to all the different ways, as we see in this paper. When you realize how quickly these things can compound, stuff gets scary fast. Back to the Anthropic article here. "It's possible that we can't build, integrate, and verify the tools that we need to understand which trend line we're actually on. Yeah, we don't know if we're getting more or less aligned if we don't have the ability to understand what's going into the training.
We do not have good intuitions for what this world would look like because our economy is currently driven by humans and human-built tools. By its nature, a world driven by fast recursive self-improvement could become dominated by the self-improving models as their capabilities fully eclipse those of humans and the model proliferates across the broader economy. It's difficult to predict what the economy looks like if human labor stops being competitive.
Even if models become fully automated and recursive, we can't predict what that would mean for most humans' daily lives. Amdahl's law applies here as well. Recursive intelligence could lead to achieving many of the benefits outlined in Machines of Loving Grace quickly in some domains. We expect that embodied intelligence like robotics might quickly follow recursive intelligence and follow a similar path of increasing returns at a decreasing cost.
More powerful intelligence might help us build things in the physical world more quickly, run more productive clinical trials of life-saving drugs, and develop novel forms of coordination. But achieving recursive improvement alone does not suggest an immediate change in how industrial production occurs, societies organize, or markets function. More intelligence can't learn what a drug does over decades of use, can't hold elections sooner than the Constitution dictates, and it can't turn a stranger into an old friend in a weekend.
Uh depends on how strong your psychosis is. For most people, the felt pace of this future will still be set by the bottlenecks, even if the laboratory upstream runs at the speed of compute. That collision where recursive intelligence building itself even faster meets the world of humans, relationships, and governance is another part of the future that we can't predict. So, what should we do about this? If it was possible to effectively slow the development of this technology to give ourselves more time to deal with its immense implications, we think that would likely be a good thing.
But if a slowdown simply lets the least cautious actors catch up technologically, it could leave everyone less safe. This is the scary reality. If America was to slow down AI development, then other less well-aligned actors would just go farther ahead, and then we're screwed. Without a global coordination mechanism, companies and governments will have to make difficult decisions about safety while under competitive and geopolitical pressures.
Yep. We believe it would be good for the world to have the option to slow or temporarily pause frontier AI development to enable societal structures and alignment research to keep up with the advance of the technology. Like like literally similar to a nuclear peace treaty across the globe. Only way this would be able to happen and people would still be quietly breaking the law, like for sure. Their example here would be the Anthropic Institute conducting research in collaboration with many others and taking action to help build the systems that a credible slow down or pause would require.
These systems would enable frontier AI developers to verify that others globally have actually stopped or slowed and that a bad actor could not use the not going to pronounce auspices, right? Regardless. The pieces of accord needed slow down to jump ahead in secret. If such systems existed, we expect that we would slow down or temporarily pause if other developers at or near the frontier also did so in a verifiable manner.
They're outright saying here if they could get others to agree to pause and we could verify that everyone paused, they would do it right now. A meaningful slow down or pause would require multiple well-resourced labs at or near the frontier in multiple countries agreeing to stop under the same conditions. What they're saying is Mistral's allowed to keep going. It would also require that each can verify the others have actually stopped.
Due to the unique characteristics of AI systems, the detectability, which is a much lower standard than verifiability, element of this arms control problem is much more challenging than with other technologies. I agree. You can't like look for nuclear particles emitting places that they shouldn't be because this is just comes down to GPUs. You could track where Nvidia GPUs are going, but even that could be kind of worked around now.
Training runs are far easier to conceal than missile silos. Their inputs are general purpose and the incentives to defect quietly is enormous because whoever continues while others pause could inherit the lead. Credible pause also specify what triggers it, what lifts it, and what adjudicates it. Who decides when it's over? None of this is necessarily impossible in principle. The world has built verification regimes for other complex technologies like the Intermediate Range Nuclear Forces Treaty.
Kind of hinted at that before. But those regimes took decades to build both the infra and the trust. We don't have that long. A unilateral pause by one lab, by contrast, is achievable immediately, but it accomplishes much less. It would change who the front runner is, but it would not create the wider deliberate process that is currently missing. I agree. This is them saying this is why we're not pausing. And while you could look at this as them justifying going against their original mission of making safe AI by rushing to the finish right now, I'm aligned with them on this.
I am not going to on them for this. I think this is reasonable. This is aligned with their goals, but I could see the conspiracy angle here where you try to say that they're doing this to justify becoming a trillion-dollar company. Like it's not a coincidence that the same range of time where they became a trillion-dollar company is also when they come out and say we're not pausing AI development despite how unsafe it might be.
Those things are aligned, but not because the trillion-dollar valuation means they can't stop. Rather, they have gotten so far that their valuation is much higher and at the same time their concerns are growing greater. They are planning on trying to jump in front of this though. In the coming months they plan to organize conversations where policymakers, researchers, civil society, and other AI companies can help answer some of the questions that this piece raises, especially around full recursive self-improvement and how to create better options for coordination and deliberation.
We'll publish what comes out of it. The window to investigate these questions together is here and people outside AI companies should be involved in this deliberation. I was skeptical going in. I had a lot of people say I needed to read this and look into it for content and I'm very thankful I listened to them because this is fascinating. The idea that they are pretty outright saying we would pause if others would, but we're scared of what happens if only we pause is wild.
I never thought I would see them be so direct with this in publications, especially now. But, yeah, like to their credit, they are holding strong on their original perspective. And this does ask some very interesting questions. What happens when AI gets self-improving? I don't know. And anybody who can confidently tell you they know probably doesn't either. Everything is changing really fast, and I am thankful that Anthropic's at least opening up the conversation of what happens if it starts to change too fast, because we don't know yet.
And we should at least start planning for the different directions things might go. I've never had a more build safety not guardrails moment than here. If we can't steer this anymore, we should at least be prepared when everything falls apart. And I hope that this video helps you start to think about these questions yourself. Let me know how y'all feel and how scared you are of a future where AI can improve itself, because I will admit I'm quite scared of this, too.
And until next time, peace nerds.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.