Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Theo - t3․gg · @t3dotgg
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in Theo - t3․gg's most watched videos.
Most replayed moment at 2:36
5.6x that video's typical replay level
Well, you have nothing to worry about cuz the first million users are free. Get yourself enterprise ready at soidiv.link/workos. I'm very excited to read into what Linear is cooking here. You know they're cooking something different cuz this is the only not dark mode page I've ever seen Linear ship. They actually took
Said at 2:29
Most replayed moment at 25:09
3.4x that video's typical replay level
just set up now. It's really nice when you do need to do things in the GUI or format the machine, that type of thing. It's time to show you guys how I actually do work using this setup. NPX T3 at nightly serve. I do have these host commands to make it work better
Said at 25:02
Most replayed moment at 2:53
29.0x that video's typical replay level
of different places in order to pull it together. Setup couldn't be easier. You click start, you add a new database, you get a connection string and now you're good to go with a real Postgress database with all the power of ClickHouse behind it. Stop compromising on your database today at soy.link/clickhouse. The best
Said at 2:45
The graph counts replays. It does not show where viewers stopped watching.
Words
7,367
Runtime
33:48
Speaking pace
218wpm
Reading time
31min
218 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Coding is solved. Bugs are not yet solved. Fix incoming. Oh boy, it's been a bit since a Boris tweet annoyed people as much as this one. I miss the days where he would just randomly post about how Cloud Code has replaced all of their engineers and everyone would get mad. But at the same time, a lot of the things he said before now kind of seem right. They were obnoxious at the time and were just very far off, especially the timeline aspects of what he said. But the reality of engineering nowadays is that if you're writing your code by hand still, you are largely
109 words, the words spoken in the first 30 seconds at 218 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 385 |
| Average words per sentence | 19.1 |
| Longest sentence | 113 words |
| Questions asked | 25 |
| Sentences containing a number | 19 |
Most used terms
Filler phrases
77 in total: like 41 · actually 27 · kind of 3 · literally 2 · uh 2 · right? 1 · you know 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
Coding is solved. Bugs are not yet solved. Fix incoming. Oh boy, it's been a bit since a Boris tweet annoyed people as much as this one. I miss the days where he would just randomly post about how Cloud Code has replaced all of their engineers and everyone would get mad. But at the same time, a lot of the things he said before now kind of seem right. They were obnoxious at the time and were just very far off, especially the timeline aspects of what he said.
But the reality of engineering nowadays is that if you're writing your code by hand still, you are largely falling behind. And I have to do a thing I am not excited to do. I have to defend what Boris said here because a lot of people are looking at this in the wrong light and pushing back in ways I don't necessarily find fair. Matt PCO's push back in particular is beautiful. Been thinking about this quote for a solid 24 hours.
It's like a raspberry ripple of a vanilla ice cream of VC funding stirred with a little pinch of turd. And the hottake you probably don't expect from me here, Matt's also right. And I want to do my best to break down what both people are saying here because they're using the same terms to say very, very different things. And if you understand where they come from and what they're actually trying to say underneath, I think you can learn a lot about where software is going and how people are applying these things in their work.
But first, I have to do my job quick, which means it's time for a sponsor break. I've been landing a lot of PRs recently. Like a truly insane amount. Before filming today, I ended up landing like eight PRs on a day where I also had a bunch of meetings. Sorry, it looks like it was actually more than that. It's crazy. And today's sponsor is one of the pieces that has helped the most here is the CIU Blacksmith. These guys have revolutionized what I'm willing to use GitHub actions for.
I was just so used to them being really slow and unreliable that I just stopped putting things on them I didn't have to. And then I started using Blacksmith and all of our times got cut by like 2x or more. There's a lot of layers to what make Blacksmith so much better than traditional GitHub actions. From the hardware that's two times faster to the cache downloads that are four times faster to the absurd 40 times faster Docker builds.
All of that by itself is enough to feel a real difference. What's even cooler is that their documentation is so good that your agents can parse it really well. I asked my agents how I can speed up CI even more and they found a bunch of ways to do more parallel work. But more importantly, they found sticky discs from Blacksmith, which lets you set up a chunk of your disc that is stored across runs, which make a bunch of things from downloading a ton of node modules to giant file transforms to your old huge git history, all way, way faster.
I look at this chart and I cry a bit because I've seen GitHub action cache downloads take way over a minute. Some of our CI spent more time downloading than running. But if you switch over to sticky discs, there's literally no download speed because it's already there. It takes 3 seconds instead of over a minute. If you want faster CI today, tell your agent to check out sidv.link/blacksmith. Okay, good to have you back.
Let's start here with Boris's tweet. This is a response to Peter's post showing a bug in the Cloud Code desktop app, which admittedly is a terrible showcase of what AI models are capable of because there is no excuse for the Cloud Code desktop app to be in as rough of a shape as it is. The team is working hard on it. They've taken my feedback surprisingly well. They were a little beat up when I first called out how bad the Cloudco desktop app was, but they realize I'm doing it in good faith.
I just want it to be better. So now they're making real improvements. Those improvements involve shipping updates all the time. And when you get an update, you can't actually read the update text because it gets truncated because a model wrote this field. And when the model wrote it, it was just reading the code and it didn't realize that the window space available for this little callout thing wasn't long enough to actually fit the text being rendered.
This is such a common mistake that I see all the time with AI generated code that frankly didn't happen a whole lot before AI generated code took over. If you were building a feature like this, you would probably make sure it renders in the browser while you're working on it so that you could see how it looks. Hypothetically speaking, you could just write the code and put up the PR having never tested it. But in a world of humans writing code, why would you do that?
Why would you waste the time of your co-workers with code that you never even saw the output of? You'd probably make sure it shows up in the browser at the absolute least. And when you did that, you would notice this edge case and fix it before you bother your team. By default, agents don't do that at all. And the result is things like this. In a world where we've learned the agents write good code and it's usually safe to merge, we get really used to doing that where we just hit the merge button if we look through the code quick and everything seems fine.
We've always trusted that original developer to test the thing they changed and make sure there are no bugs before we put it up for code review. But now we live in a world where the code review going up does not imply any effort was put into checking the thing ahead. And you end up with stuff like this. This is why Peter's post quoting the coding is solved thing is actually hilarious and rightfully so. Because while agents can write the code, they're missing so many of the things that we would expect modern software to do and we would expect more importantly for the humans writing the software to get right.
It is also absurd that we can't hot reload agents that are supposed to run 24/7. I wish it was easier to do that. I understand technically speaking why it's not. I have been working on this a little bit with T3 code, so I can empathize to the difficulty of this problem. But I think the bigger call out here is that you can't actually read the text in this. Boris's response here is less about the hot reload thing and more about the visible bug in this image, which is why his follow-up saying that this isn't a bug.
It's a UX issue that we have a fix for. This is a bug. I got to reply to this one. Boris, I'm doing my best to defend you here, but this reply is making it harder. By what logic would you say the text being cut off in all reasonable display sizes isn't a bug? I'd be really curious what your definition of a bug is because this cannot be the intended experience. In order for us to have a conversation about this, we need definitions for the words coding and for solved.
I hope we all know what is means. Coding can mean a lot of different things. It can mean telling somebody about an idea and they go build the whole thing, turning the idea into an experience. That whole end to end could be coding. Coding could be just the part where you open up your editor and you write code. Coding could be the process of turning a Jira ticket into a pull request. Coding could also include the process of reviewing that code or pulling down someone else's changes.
Coding can be a very narrow task or a very wide one depending on how you want to use the term. And let's be real, even before AI, how much of an engineer's time was actually spent writing code in their editor? If we think of coding as the endtoend aspect of what an engineer does, obviously coding is not solved. If coding is the thing I do from when I walk into the office as an engineer to when I leave, no, AI has not solved coding.
But if we think of coding as the thing that happens when I open up my text editor and I'm pressing keys on my keyboard, if we narrow that definition down to be that small, where coding is the act of typing code on my keyboard inside of an editor, yes, absolutely coding is solved. I'd even widen a little bit and say if you include things like finding the right Jira ticket, going through the stuff that users are doing, turning the data from the experiences users have into the code that you want to write, that whole experience from like rough idea to code that works.
That all can be done with the AI now. And when it comes to actually generating code that functions, AI has solved that. But engineering isn't just coding. Going to do a silly example. Please hear me out. This is the password game by Neil Fun. This is meant to mock how hard it is to sign up to most websites nowadays because their password requirements are absurd. So, every time I fix what is missing, I have to make additional changes.
Now, it says the digits of my password have to add up to 25. Cool. There we got that. Password must include a month of the year. April must include one of our sponsors. Hilarious. Okay, so why am I showing this? This is obviously a terrible experience. If this was the password system in a real website, we would make fun of it endlessly because it's bad. We don't want to use this. But was Neil coding when he built this?
Obviously, yes. When Neil made this game, he coded it. He had the idea. He opened an editor. He wrote code to make this thing exist. And now it exists and we can experience it. He did intend for it to be this rough, buggy, messy experience because that's the point is making fun of these types of experiences. So he did it with intent, but he did code this. And I hope we all agree that the process of somebody like Neil having this idea and then putting it out there, that coding is at the very least involved heavily in this process.
So now I have to ask, if somebody built something this bad accidentally where they were told to implement a password field in their website or web app and when they did that they made it really rough and miserable like this, were they coding? I would hope we can also agree the answer there is yes. That just because the code didn't come out good or useful based on the actual goal that a user would have using it doesn't mean they weren't coding.
I hope we can also agree there's a difference if this type of horrible experience happens because there wasn't enough detail in the request made to the developer where if I told them make a password system for our site and then they went and whipped this up because they didn't know what a good or bad password like experience was that is different from if I write a bunch of tickets that say we have all of these crazy rules make sure that you apply all of them and the developer creates something with all of that bad decision-m applied to it.
These are very different. Bad code happening because of no detail given in the planning process versus bad code happening because of bad details in that same planning process. Missing details can lead to code that doesn't behave as expected. And bad details can also lead to code that doesn't behave as expected. But that's the input side. Bad info going in, whether it's missing stuff or is just bad with the details it has, going into the code box, results in bad code.
But there's another thing that can often result in bad code, which is the other side. Not that the code comes out bad, but because you don't have a system that allows you to verify the code on the other side. Let's look at it like this. You have these different qualities of plan. You have the engineering cycle. It could be one person. It could be a team. It could be three engineers doing a sprint. It could be whatever.
And when you put in a good plan, we hope that out of this comes good software. There are plenty of ways that this can change. Obviously, if the quality of the engineers isn't great, if the good plan gets lost along the way, if somebody rewrites the plan and goes off course, but generally speaking, if you have good engineers and a good plan, we hope that the result would be good software. If you have a bad plan, the most that we can hope the engineers do is give feedback on that bad plan and say, "Hey, I'm not sure if this is the right idea." But if I've learned anything from my time as an engineer, it's that people outside of engineering often don't like when you come in and say that their plan for how to change the software isn't great.
And I understand it's their job on the product team to figure out what we should be doing, not the engineer's job. It's our job to do it. But I would then ask you, is a truly exceptional coder, a person whose job is to turn ideas into code, are they going to push back on a bad plan or would they just implement that bad plan because they're good at coding, not making good software? Because those are different things. a bad plan going through great coders that aren't there to push back on plans.
They're there to execute those plans. They're going to make bad software because the thing going in is bad. The thing going out is obviously also going to be bad. And then there is the empty or bare plan where there just isn't enough detail in the plan. This has the widest range of potential output. Like if the plan is good and the engineers are good, we can say the range of success here would be like 8 to 10 obviously out of 10.
Where with the bad plan, I would argue that the software coming out will also probably be bad as low as like a one to a four out of 10. Regardless of how good the engineers are, this just kind of affects what range you're falling in there. And then with an empty or bare plan, the range is massive. I would say as far as 1 to 10, you have no idea. It depends on a lot. It depends on how complex the thing is that is being requested.
Depends on how well the engineer understands all of this as well as how well they understand the needs of their users. how complex the software is, how complex the request is. There's so many layers to this that the range, I would say, could be anything. If you give great engineers an empty plan that is a very, very vague idea of what you want, and they're excited to just write code and put up a poll request, the likelihood that this is good is random, almost probably not looking great if I'm being real.
So, what do we do to fix this? How do we handle this as an industry? Obviously, step one is to write better plans, but if we can't control the quality of plans at all times, we need to make sure we have other solutions. We also need to make sure that the code that comes out is matching the needs of the plan because we don't know if the engineers in this box are good or not either. So, how can we improve our outcomes in the box here?
I would argue we've had a solution for this for a long time now. It's the verification layer. things like QA, things like staging environments you can test in, things like slow rollouts that will make sure that bad code doesn't hit users and if it does that it minimizes the impact of that bad code. Good verification steps will allow for your engineers to verify the results against the plan and perhaps in some cases even notice that once you implement the plan, the experience isn't great.
If I have a plan that looks good, I can't tell it's a bad plan because it all seems fine and then I build it and I put up the code, it's unlikely that I know the actual experience is bad until I try it. Trying it could be spinning it up on my machine. Trying it could be testing the preview build when it happens. Trying it could be your agents going through the loop with it and showing you a video of it when it's done.
But you need something here to prevent the bad plans and even to an extent the good plans from having things sneak in that are not a good experience. Good plans and good engineers should result in good experiences, but I would argue they only do when the good engineer has a system to verify the changes are actually good. But if you have a bad plan or bad engineers, then you need this even more in order to verify the quality of those outputs.
Hopefully, you could already see what I'm about to do. Let's replace this team of engineers with who is really doing the work today. Swarms of agents. You take your plan, you give it to Claude, you give it to Codeex, it spins up a bunch of sub agents to execute all the parts, and then it puts up a poll request. If it hasn't actually run the code to check, it might look fine, it might seem fine, but you don't know if it is fine until someone checks it.
Usually, that'll be a human. And sadly enough, the human is often the one using the software in the end. It's not one of the humans who work on the product. And that's exactly what happened here. Do you legitimately think after this code was written by an agent that any human or agent actually checked what it did in the app? Obviously no. Because this would be a very quick thing you see and you're like, "Oh yeah, that is wrong.
I should fix it." But no one did because no one saw it. And by the time it has merged, the friction to unmerge it or to fix it is higher than the friction to build and merge it in the first place. To go back to our diagram here, let's say that after the verification loop, you've determined that the software is good. At the very least, based on your understanding that the problem you're trying to solve has been solved and the thing that comes out in all obvious ways seems fine.
And I would expect that the result is good software, probably similar range. Maybe we could even bump from that 8 out of 10 to like 9 out of 10 or 10 out of 10 range. What happens if the result isn't good? Well, ideally that bad result will send you back up to the planning stage so that you can adjust the plan based on what you've learned from the verification. Maybe the bug is just something silly in the implementation and you can skip straight back to the agent swarm or the engineers working on it to have them fix whatever isn't right.
The point I'm trying to make here is that if your verification system is good, then before the code ever goes to users, it should be able to identify the problems and fix them as they are working. And right now, agents don't do this at the very least by default. And this can be caused by a lot of things, many of which are actually human error. I cannot tell you how many code bases I've worked in where just spinning it up on my machine to see if my changes worked was more complex than making the change in the first place.
And it ends up being easier to just file the pull request and wait for the preview build to come up than it is to try and test it on my machine. If it is too hard for you to spin up and test, it is way too hard for your agents to do the same. You need to make it easier to actually verify your code bases. And I find the companies struggling the most with this are the ones who are building in a way that is hard to verify.
Funny enough, Anthropic is one of the most guilty of this. Since Claude Code Desktop is a full desktopbacked Electron app, in order to run QA and verify it, they need to give every thread its own graphical VM fully backed with a real instance of an operating system that is supported with Cloud Code Desktop so that it can run the build, spin up that Cloud Code Desktop instance in an environment where Claude can control it and check it and prod at it.
And when you combine the fact that that is expensive and difficult and far from trivial to set up with the additional fact that cloud code models are just so much worse at computer use than models from other places, you end up without that good quality verification loop that is necessary to prove this software works. It's silly to put it this way, but I think a significant portion of why the Cloud Code desktop app isn't improving meaningfully in comparison to the Cloud Code web app, which is improving constantly, simply comes down to how easy it is to set up Cloud Code web for testing versus Claude Code Desktop.
Since the Cloud website is built on, you know, the web, it's a lot easier to set up a loop where your agents can test it. You could even have one given machine where your agents have eight tabs open for eight different dev builds checking different things, and it's totally fine. with Cloud Code desktop. Good luck launching eight separate Claude Code instances as a desktop Electron app on the same computer that I can poke and prod through without getting confused along the way.
It's not going to happen. But this would also be hard for a human. If I was trying to review two different PRs and test their changes at the same time on my computer, it wouldn't be pleasant. I know cuz I've built Electron apps for a decade. If you have two things you want to check at the same time, you're kind of just screwed. And this means you have to build a system to make it easier. Again, if it would be hard for the human to verify, it is probably impossible for the agent to do it.
And quadcode desktop is a golden example of what happens when you don't have a good enough loop for catching these types of things by actually testing the changes that you make. So going back to this diagram, my question for you would be what parts of this are coding? Do you think coding is narrowly scoped to this part here where plans are transformed into code? the act of typing on the keyboard in the editor or the agent transforming these ideas into source code that you can actually run and use or do you think it's wider?
Do you think coding also includes the planning side actually figuring out what you want to do? Do you think it includes the verification side making sure the changes are valid or not? Do you think it includes the full loop where that verification layer goes back into the coding part? I think everyone defines this a bit differently, which is why the conversation has gotten as chaotic as it has. Because for some people, coding is just this section in the middle, the agent swarm part where you're writing the code.
And for others, coding is everything you do going from idea to functioning software. I can hear the argument either way. I personally don't care that much. But at the absolute least, I would say that the whole process here from start to end is software engineering. And if anybody tries to say that software engineering is solved because agents write code, okay, in this middle section, then both Prime and I will be equally mad at them because software engineering isn't solved.
Software engineering isn't just this little piece here where you make code. Software engineering isn't even just this square. It is far wider and forces you to think about the architecture of the systems you're building so that this whole loop can happen more effectively hundreds of times a day. The engineering isn't just what you do inside of this box. It's the platform you build for the boxes to sit on and stack up and continue to grow.
But Boris didn't say software engineering. He said coding. And that's historically what he has said. If Boris had said software engineering is solved, I would be making a very different video right now. But he didn't. Bugs are part of software engineering. The same way that building good verification loops is part of software engineering. Architecting systems in a way that you can make changes confidently is part of software engineering.
These aren't just coding tasks. These are engineering work and that's why that word exists because it's not just the task of being a code monkey sitting at a computer typing letters on the keyboard. And if I'm being really frank here, agents need a lot of help with this part. The same way that an incredible engineer that you just hired would need help with these things in order to work in your codebase. If a new engineer showed up and they were assigned this change and they couldn't even get the test build running on their machine, much less three of them so they can test your changes too, they're just as screwed as the agents are.
And I think a lot of us have code bases where a great engineer would be unhappy working in them and wouldn't have the resources they need to verify their stuff. Now imagine that person has been blinded that you literally force them to work with their eyes closed. How good of software do you think they're going to make? Probably still decent here and there if the changes don't require them to see what they're working on, but it's a lot better if you give them what they need to verify those changes.
So when I'm hearing Boris say coding is solved, bugs are not yet solved. What I'm hearing him say is that this part here agents have gotten incredible at. If you give them a wellsp spec change, they will make that change. The change might not be what you actually want, but if you tell them what to do, they will do it 99% of the time. And if they get it wrong, you can tell them what they got wrong, and they'll fix it 95% of the time.
But the next part, which is finding and preventing bugs autonomously, ends up being significantly more difficult. Not just because the agents are stupid and don't know how, but because our code bases are stupid and don't set them up for it, too. We need improvements in our system and architecture in our apps in order to make them easier to test and verify. We need better tools from the labs and people building AI development tech in order to make it easier for the agents to actually access those systems that we're building and test things in verifiable, repeatable ways.
We need better systems for our agents to show us that they actually checked their changes and to make that more normalized. Hell, we need a way for GitHub to allow our agents to upload an image or a video proving their work without having to commit it to the repo, which has caused all sorts of disasters across a lot of projects. Fun fact, there's a new GitHub CLI update where they did actually finally add this, which I think is a huge stepping stone in getting all of this right.
But that's the key. We need to figure out both what we can do in our code bases to give agents more likelihood of verifying things well themselves as well as expect the providers we rely on whether that's anthropic and open AAI or if it's GitHub and Google Chrome whatever the layers are we need them to be exposed in ways where our agents can verify their work and prove that they have done such too and even then I've seen a few too many times where an agent shared a video of it proving the work it did is good just for the video to be obvious slop where the thing didn't work at all.
I cannot tell you how many times I've had to say, "Wait, the video you shared shows this not working. What did you see?" And then it realizes it got it wrong and it goes and fixes it. Another real hot take I have on this before we moved to Matt's post. I think if you have the bug speced out properly, like a detailed enough bug report that describes what is wrong and what the fix should look like and then add even the most minimal verification system there, the vast majority of frontier level agents and models can absolutely solve the bugs, too.
They're not going to ship software without bugs. They're not going to preemptively find these bugs and fix them as often as we'd want. They're never going to ship bug-free software that's just not really possible. There will always be edges that you didn't plan for in your systems and in your verification. But if you make it easier and easier to identify these bugs and the models get smarter and smarter at fixing them, then yes, to some extent, bugs will start to be solved as well.
Imagine a world where the agent starts working on the code. It makes the changes. It wants to test them. It realizes it can't because the repo isn't set up properly or its tooling isn't set up properly. So, it then goes and builds all the pieces it's missing in order to verify its own changes. I can't tell you how much time I've spent with T3 code setting up things so that the agent can actually spin up and test work as well as expose it to me.
When I'm working remotely and I have my box set up over Tailscale on another machine and I want to see what changes it made in its dev server, I can't click a local host link when I'm on an entirely different network. I need it hosted through tail scale. So I actually made a lot of changes like thousands of lines of code changed in T3 code to make it easier to run the dev server over a tail scale environment while also pulling in a readonly snapshot of your existing data to make it a more useful test.
These are the types of things I have done both to make it easier for me to verify the changes the agents made as well as for the agents to verify their own work too. In the future, agents can notice that they need these things and unblock themselves and make suggestions or even full-on create a new poll request just to set up your repo for them to catch bugs themselves, which will massively reduce the number of bugs that they are shipping.
But until then, it's our responsibility and we need to build the systems the agents need to verify the bugs and we need to let the agents know when the bugs happen. And most importantly, we need to check the results, not necessarily the code. Because let's be real here, nobody would have caught this bug in code review because the code looks fine. If it didn't, the agents would have solved it. The era of finding bugs by just reading the code has, let's be real, ended.
Any bug you notice from just reading code can almost certainly be noticed by the agents, too, and ideally noticed within their own development loops. But you still can notice a lot from reading the code, like architectural failures, not using something that you should have been, using something you shouldn't have been, touching a place that shouldn't be touched for these changes, writing tests that aren't necessary, not writing tests that are necessary.
There's lots of things you can see in the code, but those aren't bugs in the traditional sense. And the bugs can absolutely be seen by an agent if the bugs are visible in the code itself. But sadly, a lot of bugs aren't, like the one we have here. What that means is we need to be actually using the code or agents, right? We need to pull it down on our machines and test it and check it and make sure it solves the problem it's intended to at the absolute least we need the agent to show us that it did that itself.
And until the agents can do that whole loop themselves, bugs are definitely not solved. So hopefully you now understand my stance on this and what I think Boris was saying. So let's wrap up by going through what Matt said as well as Boris's responses. Matt saying that this post reads like a vanilla ice cream of VC funding stirred with a little pinch of turd. I think this is a totally reasonable read of what Boris said because he said it in a in a way meant to ruffle feathers.
That's what the whole coding is solved thing does. Especially if you have different definitions of coding. If one of these people is using coding as a term for what they're doing on the keyboard and the other is using it for the entire software development life cycle, they can't agree because they're arguing the same term with different meanings. But Boris didn't offer any clarity in his post here. So, it's a totally reasonable way to read it if you use coding as a way of describing the whole life cycle.
Boris replied, "TBH, this is a good debate to have. Here's what the timeline of the near past and future looks like to him. One, models are able to code better than he can. Two, models are able to do coding adjacent engineering work better than he can, like debugging, profiling, optimizing, bug fixing, system design, abstraction design, UI design, idea generation, etc. And eventually models are able to do most things that can be done on a computer better than most people.
For the types of coding work that Boris does, Claude has achieved the first, which is that models are coding better than he does and some, but not all of two, which is that they can do coding adjacent things like system design, bug fixing, optimizing, profiling, debugging, etc. But it's only starting to show early signs of three, which is doing most things doable on computers better than people. This is not the case for all of coding yet.
If you're Andarus Hellsburg and are a worldclass expert in compilers and type systems, Claude might not be at one yet. Boris is betting that it will get there quite soon while at the same time pushing further into two to three territory for the rest of the world. I would argue that we are actually surprisingly far into one for those types of god engineers. I know that for example Ryan Carneato has been going more and more deep into fable and realizing that he can use it not just to like make small changes to the framework he built solid.js s but to actually talk out complex ideas and test theories that would have been too timeconuming for him to do.
So I am seeing more and more of like the top upper echelon best engineers in the world realizing that they can utilize these agents to do a lot of coding work not like as incredible as them but close enough to it with a lot more parallelism that lets them test things that they would not have been able to otherwise. Back to what Boris said here. You can think of model capabilities as capturing the distribution of human ability.
I consider myself an average programmer. So Claude has already surpassed me. The same is not yet true for everyone. Back in November last year, Boris fully stopped writing code by hand. Though he continues to code using agents every day. That's the point where for him and many others around him, it felt fair to say that coding is solved. But engineering is more than coding. Claude's code is not perfect. It has bugs and inefficiencies.
Bod can code, but it cannot yet do everything that goes into engineering. I swear I did not read this before deciding to make this video or filming the section before. It's funny that he makes the same distinction, but again, I had a feeling that's what he meant. Opus 48 was the first model that felt like it could be better than he was. Babel now routinely finds optimizations and debugs issues that would have not been able to be found by Boris himself.
And with each generation, the code the model produces continues to improve and the model's able to do more of the non-coding parts of engineering. Similar to how I had pushed back on Boris earlier with his definition of bug, because he said that it wasn't wasn't a bug, it was a UX issue. Matt's asking for his definition of coding. He cites John Asterhout's definition of tactical versus strategic programming. Tactical is the on the ground day-to-day aspects of coding, and strategic is the long-term stuff like codebased health, architectural design, making the right decisions.
Matt agrees that AI's largely solved tactical programming, but he has seen no evidence that it can think strategically. Asterout talks about some employees as tactical tornadoes, able to churn out astonishing amounts of work with zero high for the future, and that's how agents feel to him right now. I largely agree. I don't think agents will preemptively, proactively do good with strategic planning of codebase health and architecture over time, but I do think they can be useful resources during that workflow.
If you ask agents to look into these things and give you feedback on how your architecture looks or what changes you're planning on making and you can even have it investigate theories to prove ideas you have out too. There's a lot of power to be had with agents helping you as consultants when you plan your strategy. And then once you've done all of the strategizing and you have a plan you're confident in and you give your agents the things they need to verify their outputs, all of a sudden they can go execute a crazy bold plan in time that makes no sense at all.
It's truly unbelievable how fast these agents can work if you set them up for success, which uh spoiler has almost always been the case with engineering. If you set up your systems in a way where everyone at the company could easily pull down someone else's changes and verify that they fix the issues they have, where it's easy to make those changes and find where in the code base is relevant to the stuff that you're doing, making it easy to actually put the code up for review and to get it approved and shipped and rolled back if it does break things.
All of these decisions you make around how the day-to-day works in your codebase would have made engineers more effective when you brought them in. And you can bet your ass they're going to make the agents more effective, too. It's actually really funny reading all of this because I swear I had the whole video planned before I did and they were doing this back and forth right now. But it really does line up with what I was saying there where Boris splits coding and programming up as the act of writing code and engineering as everything else in addition to that.
Tactical and strategic is another reasonable split. He does see some parts of strategic coding being automated like when they're starting to automatically maintain apps at Anthropic. Recently a number of the customers are too. We haven't bit this bullet quite yet at T3 Code. I'll still screenshot the tweet and go post it into T3 Code myself. I won't have the agents propose fixes based on what people are tweeting yet.
Matt calls out that he agrees partially on the strategic coding being automated thing, saying that people are building pipelines like bug report to repro to fix feature request to prototype or cron to codebase architecture RFC to refactor. Managing these pipelines is very much a strategic concern, one that a human needs to apply their judgment to. Absolutely agree. In order to make your agents act more strategic and act more like engineers, you have to do the engineering upfront yourself.
The agents can help you as you do it, but if you want the agents to be able to do these awesome things, you have to set them up for success. Man, I thought this was going to be a quick 15-minute video, and I'm looking at the timer and realizing it wasn't. Thankfully, I filmed this offline for once. I know usually my videos are done on stream, but uh cast. Okay, I didn't feel like going live today and just wanted to talk about the stuff that was interesting to me.
And I hope you enjoyed it. This is quite a fun deep dive into all these things that I've already been thinking about and I love the excuse that both Boris and Matt gave me to talk more about it. I'm curious how y'all feel. Am I being crazy saying that engineering is still unsolved and coding is solved? Am I going too far by saying coding is solved? Or am I not going far enough by saying engineering is still largely a human task?
Let me know how you feel and where you guys are at in your real world work.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.