Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Theo - t3․gg · @t3dotgg
Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
16:133.8x the video's typical replay level
genuinely working deliberately narrow no check transpiler from a five file corpus, but everything else is code that hasn't been proven and the tests are currently failing. Currently, I have a bunch of work that was stuck in a stash, so I'm going to clean that up and figure out how far it actually is of the Obviously,
Said at 16:05
Most replayed moment #2
12:493.4x the video's typical replay level
worried about how good this model is at front end, I'm not going to say it's perfect, but when it's steered the right way, you can get good front ends out of it. I think this looks pretty good and most of this was done with 5-6. So, yeah, take that as you will. I also had it use sub-agents to go through all of the T3
Said at 12:41
Most replayed moment #3
3:553.3x the video's typical replay level
A few more things I want to clarify before we dive too deep. OpenAI has not paid me for any of this coverage. Normally, I don't do this level of usage when they give me early access. I do maybe a hundred or two hundred dollars of inference, which is covered under my paid two hundred dollar plan anyways, which I'm
Said at 3:48
The graph counts replays. It does not show where viewers stopped watching.
Words
6,025
Runtime
26:09
Speaking pace
230wpm
Reading time
25min
230 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
GPT 5.6 is finally here, and I don't think I've ever been this excited to talk about a model. There's a reason why, though. It's not cuz it's the best thing ever. It is really great, but the reason is cuz I've had access for a while now, since before Fable even first dropped. I've been waiting to talk about this model for so goddamn long that I've been putting out less videos cuz I've been trying to clean up my computer to hide the fact that I've used it so much. And when I say so much, I mean it. Depending on how you estimate my usage, I did between $180,000 and $240,000 of
115 words, the words spoken in the first 30 seconds at 230 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 323 |
| Average words per sentence | 18.7 |
| Longest sentence | 72 words |
| Questions asked | 4 |
| Sentences containing a number | 56 |
Most used terms
Filler phrases
66 in total: like 30 · actually 27 · kind of 3 · uh 3 · right? 2 · I mean 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
GPT 5.6 is finally here, and I don't think I've ever been this excited to talk about a model. There's a reason why, though. It's not cuz it's the best thing ever. It is really great, but the reason is cuz I've had access for a while now, since before Fable even first dropped. I've been waiting to talk about this model for so goddamn long that I've been putting out less videos cuz I've been trying to clean up my computer to hide the fact that I've used it so much.
And when I say so much, I mean it. Depending on how you estimate my usage, I did between $180,000 and $240,000 of inference in my month and a half-ish of access. And I got an absurd amount done in that time. It took a bit for me to digest what made 5.6 so special, but once I did, I started pushing it further than I've ever pushed a model. I did more inference in the past month than I had done in my life up until that point.
And I have so many thoughts to share now that I've done this. I need to be realistic, though. If I share everything I have to say about this model, the video is going to be hours upon hours long. So, instead, I'm going to break this up into multiple different parts. And this one's particularly special. The reason this one's special is cuz I haven't read any of the coverage on this model yet. I have no idea what people think about it, how they're using it, or what they expect from a video like this.
I'm just going off the top. I'm sharing how I used the model and how I spent such an absurd amount of money on this model to show you guys what it's actually like, or at least what it was like for me. What I'm trying to say here is don't think of this video as a review. I'll have a proper 5.6 review coming up really soon after this. Right now, I just want to show how I used the model and the cool things I built with it cuz turns out you can do a lot with $200,000 of tokens.
What that means is I won't be talking about Luna and Terra cuz I didn't have access during the testing. I won't be comparing with Fable. That has its own dedicated video. I won't talk about ultra or max reasoning levels because one of those is a reasoning level, the others are new feature, neither of which are things I had access to. I won't talk about the death of Codex even though, believe me, I want to talk about this.
I have a lot of things to say. I won't talk about benchmarks or sentiment cuz again, I haven't seen any of them other than what was posted with that initial announcement, and those benches didn't really make sense to me. I won't talk about pricing or usage limits, because again, I didn't have any of that when I was testing the model, and I won't talk about the fast version, because it didn't exist. My fast version was 1.5x faster, not the 750 TPS that's being promised by Cerebras.
I don't even know if that's out yet. Again, I haven't checked anything today. What I have checked out though, is today's sponsor. I'm going to be real with you guys, I was super, super skeptical of computer use and browser use when I first saw it popping up. It just didn't make sense to me. Why would an agent be faster than a human clicking buttons, if it has to screenshot, manually move the mouse to a specific coordinate, click, and then wait, and hope it gets things right?
And my early experiences with computer use stuff showed just how bad it was, and I moved on. But then something important happened. GPT-5.6. There's a lot of things that make this model special, but the one that impressed me the most immediately was how much better it was at using a computer, in particular using browsers. Like you can put a task like, set this up in a dashboard, make this change in Notion, and just navigate the web for me, and it would.
But then I had to dedicate my whole computer to letting CodeX use it. Well, I did, before I set it browser-based. Browser-based built the browser for your agents. They allow for programmatic access to the entire web through their Puppeteer layer that they are hosting for you. You no longer have to deal with Playwright yourself. You don't have to set up a browser or a real computer, you just give an API to your agents, and now they can do whatever they need to on the web.
And this isn't just some small startup guys. Everyone from Microsoft to Lovable to Amplitude is already relying on Browser-based for running their browsers. Even DeepMind has been using them recently, which is just unbelievable. But setting up browsers for your agents isn't easy. And in a world where 85% of the web's APIs are blocked behind UI, a browser becomes essential for doing real work. Give your agents the entire web at soda.link/browserbase.
A few more things I want to clarify before we dive too deep. OpenAI has not paid me for any of this coverage. Normally, I don't do this level of usage when they give me early access. I do maybe a hundred or two hundred dollars of inference, which is covered under my paid two hundred dollar plan anyways, which I'm still paying full price for. In fact, I have a bad feeling I might have to open up another one of those accounts soon.
I also don't think that my two hundred thousand dollars of inference is a realistic use case. I don't think anyone should be doing this every month. That was more me pushing the limits, experimenting, and seeing what I could get away with to an extent because they've been very tolerant of my OpenAI has also been too tolerant to the point where others were leaking the fact that they had early access without permission from OpenAI, and they didn't do anything about it.
So, I could have shared my thoughts earlier, but instead I chose to do what I thought I should do, which is wait until the model's officially out for everyone to access. So again, they're not seeing this. In fact, there's a good chance you're watching this before any OpenAI employee has because I just want to share my honest thoughts and experience using this model. I'm also going to do my best to not talk about Fable too much in this video.
I'm going to do a dedicated comparison later, so pardon me for ignoring the existence of Fable for a bit. If you do want to hear all of these other things like my proper 5.6 review comparison with Fable, day-to-day usage tips, trying to maximize your utilization, all of the things that make you really good at using this new model without having to spend the insane money that I spent on it, make sure you're subscribed and that you hit that bell because I have a lot of videos coming out.
I have so many things I've wanted to talk about with this model that I haven't been able to, and I'm finally unleashed. There's going to be a lot of content the next few days, and you're not going to want to miss it. I also didn't test the model in ChatGPT or 5.6 Pro at all. So, if you're curious about those, I'll talk about it in the future. As I was saying, so let's go through all of the things I did with this model.
I worked across 67, nice, projects, and I did a lot. Some of them I did not too many prompts in. I just set off a goal and hoped for the best. Others I put a lot of back and forth effort into. If you were confused why I had so many different machines on my network recently, this is a big part of why. I wanted to push this model to its limits and really take advantage of the generous near unlimited usage they were giving me.
And I did. I pushed it really, really hard. It's also worth noting that I set up my fleet using 5.6, their SSH computer use and a centralized config that I had it fan out to all my machines. It was actually really nice to work with. And that's one of my favorite things about this model. Not only is it much better at understanding intent than 5.5 was, it also just grabs onto your task and doesn't let go until it's finished.
This was one of my biggest issues with 5.5. It almost felt like it would get lost and then stop. Or if anything bad got into its context, you'd have to go reset and make a new thread. 5.6 I did not have that experience with it all. I just will let one thread rip forever and it's fine. Even something like {slash}goal went from holy disinteresting to me to actually pretty exciting cuz the model can manage that and can go for a lot longer without issues.
That said, I didn't find myself reaching for {slash}goal too often because if the model needed to run for more than a few hours, it just could. I would tell it to go do a long-running task and it will hold on and go for hours upon hours. I've had tasks go for 20 plus hours without setting a {slash}goal at all. I just told it to do a big thing and it went and it did the big thing. I guess it's time to go over all the work I did then, right?
I'll start with Lakebed cuz, I'll be frank, I really want to get this project out. Apparently, if I was paying fast mode prices, the work I did would have been about $30,000 in tokens, which sounds insane. But remember, you don't have to use it in fast mode, which cuts the price by almost 2x. And you already get around 14 grand of inference right now on the $200 a month plan outside of resets using Codex. So like, my usage here is actually kind of contained within a normal sub.
And that's kind of crazy cuz I set up 5.6 on some impossible tasks and it went way further than I ever would have expected. One of the biggest things I did was move from giant monolithic JavaScript files that 55 wrote into a properly broken down and managed TypeScript project. This was a huge change that required a lot of verification and just hands-on work, and it did all of it without much issue. It also helped me build a whole new CI pipeline and preview and release workflows that I'm still using to this day.
I'll probably use indefinitely. It created artifact storage, railway buckets for actually managing files that users can upload, fixed a bunch of core stuff, and also handled a bunch of chaos with refreshing deployments with Postgres events. I built CLI login. This was way more complex than it seems to imply here. Building a relationship between my production server and the CLI that makes sense that allows you to authenticate.
Not trivial, and it did pretty much all of that by itself with just a little bit of guidance on what I was expecting UX-wise. It implemented whitelist access, which is a ton of work. Probably the single biggest thing I had it do in Lakebed because it had to centralize off into its own system to do that. I had it fork Shrew my off service and build that all in, and it works great now. I cannot wait to launch the current staging version of Lakebed as the production version cuz the vast majority of the changes were written by 56.
I hardened the capsule database layer to make the reactivity way more efficient and reliable. I designed and repeatedly implemented and reviewed the developer files contract, opaque identifiers, mine policy, all all the file object storage stuff cuz I wanted to make sure that was all really, really good and reliable. The first-party off layer, which I mentioned before I did for both the whitelisting and for the CLI. I had it research different isolate options for the runtime that runs the code that people ship on Lakebed cuz remember, it's a cloud platform to build real apps on.
I wanted to make sure if I was to build my own isolate layer that we did it right, and it explored a ton of options and gave me really good HTML plans I could read similar to what we're reading right now. It helped me model pricing and quotas, which I'll talk a lot more about my thoughts on what it did here because spoiler, other models did this in a I much, much preferred. It recovered a ton of PRs for me. I showcased a few loops I did in the past and to show how powerful they were when I had the model spinning up threads to review PRs and whatnot.
When I actually did that, all of it was 5.6. I had to go through and manually swap them all to look like I was using 5.5. I'm sorry I did that. I did not realize just how much worse 5.5 was at sub-agent stuff than 5.6. 5.6 blasted through that. I tried the same prompts later on with 5.5 and it just got confused and lost. It also did a bunch of auditing for my open-source launch for Lakebed, which I'm very, very excited about.
It landed a ton of code, dozens of PRs, and really helped to turn Lakebed from a weird small side project into a thing I'm excited to ship. And I'm so excited to ship it. You guys have no idea. Then, of course, we have T3 Code, the thing that I did a lot of my work with. T3 Code is the project Julius and I have been working on non-stop to make it easier to manage your agents on whatever machine they're on from a unified layer, whether that is the desktop app, the web app, or a phone app, which is actually very, very far along, too.
Early access for the mobile app coming sooner than you think. My work on T3 Code showcased just how thorough and thoughtful 5.6 can be, in particular with mobile. I've had so many problems with different models on mobile, even modern, really high-end ones that I'm trying to not just name-drop constantly. 5.6 is so much better at mobile that I realized I needed to push it harder, so I did. I had it rebuild the entire React Native app using AppKit and Swift, as well as SwiftUI separately.
So, that's two native rewrites from scratch. And I was blown away with both. They were actually fully complete, end-to-end, functioning with all the same features that existed in the React Native version. And it did that in like two to four hours for each of them. Unbelievable. I thought it would take a day and make something entirely broken. But when you combine the capability it has with mobile with its computer use capability, where I can actually spin up the simulator and check things in it, you can get full end-to-end insane projects like this functioning.
I had a bunch of other work I did with 5-6 in T3 code, but I really want to lean into this rewrite angle because I was blown away with how much I could rebuild from scratch using the model. I've a lot more to say about these types of bold rewrites and believe me, we'll be back to that. First, I want to go over the other things I built with T3 code in 5-6. I built a full end-to-end computer use loop into T3 code so that you could do this type of iteration with mobile and computer use with 5-6 in T3 code directly.
Ended up not even putting up the PR because Julius got it working even better himself, again using the same model. We have audited a collapse worked for history, which was very, very nice because we had too much stuff going on in the history. Audited the Codex app server primitives in order to try and make sub-agents work better. I even got a PR up that did that. I don't know what we'll do with it. Designed a route back top bar computer switcher as well cuz I wanted to try and manage all my different computers in T3 code differently.
Didn't end up shipping this, but it was insane how far it got without anything beyond a vague prompt of what I was thinking of. Code thing is also T3 code. It just didn't split correctly when I asked it to combine all these projects. I rebuilt the marketing site and got it to the stage it's in today using the model. So, if you're worried about how good this model is at front end, I'm not going to say it's perfect, but when it's steered the right way, you can get good front ends out of it.
I think this looks pretty good and most of this was done with 5-6. So, yeah, take that as you will. I also had it use sub-agents to go through all of the T3 code PRs and help me orchestrate all of them to figure out what to prioritize and merge and not merge. I ended up liking the results here so much that I built an automation with Hermes to do this every single day. It has helped me stay on top of PRs better than I ever have in my life.
Here is the mobile native version. As I mentioned before, I rebuilt the whole thing in Swift UI and it succeeded. Here is it talking about that. I also just had it do a bunch of profiling on the web app to find any performance issues that are worth addressing. Okay, I need to talk about the rewrite thing because it is just so cool. The combination of the model working well for these types of rewrites and the inspiration I had from a certain bun being rewritten in a certain rust with a certain fable had me inspired to try doing similar rewrites.
So I did. I started by trying to rewrite Hermes agent in rust with complete parody and ended up getting pretty damn far. One of the interesting things I did when I was working on this is I had it archive all of the stuff I had done with my Hermes agent as well as Ben with his. So I had all of our history and then I had 56 go through it and figure out all of the features that we need in Hermes in order to use a port of it and it came up with the 50 or so percent of features that we actually used and went and implemented almost all of them.
I actually got it spun up in its own separate channel and responding creating threads using all of its skills calling the model and doing the things I expected to do with a goal I ran for about a day. It works. A full from the ground up rewrite in rust that uses like 15 megabytes of RAM instead of the giant Python version that is not the most performant thing but that's not the goal of Hermes. I have it on a dedicated Mac Mini anyways, who cares.
I was just curious what it would look like to do this and make a very small memory safe reliable minimal Hermes agent style thing that I could put in a small VM and throw wherever and it went way better than I expected. If I was to really sit there and work on it further, I bet I could finish this without too much additional time. Not anywhere near feature complete against Hermes agent but to a point where I would be able to use it every day without issue.
To be fair, that would have been about 13 grand on fast mode probably closer to five to seven if it wasn't. Still pretty insane. Not as insane as this port though and this has been very hard for me to not talk about because I had this one going for multiple days. I started rebuilding the TypeScript Go port in rust. Not because I thought it would work or I thought it would be useful, just cuz I was curious how it would go.
I was actually able to get a 100% working TypeScript compiler, specifically the transpiler that turns the TypeScript into JavaScript, all written in Rust that is up to 18 times faster than the equivalent Go version. But, I didn't get the type checking very far. I tried and I had a lot of tests passing, more than I would have expected. The code base that it built is pretty wild. It ended up being almost 200,000 lines of Rust and a lot of tests passing, although admittedly, according to Fable, which I just had analyzed the code, it's a broad prototype with a tiny verified slice and it's far from shippable.
It's 15 to 20% of the way to a usable tool, but only 5% of the way to a proper TS Go replacement despite over 195,000 lines of Rust across the 29 crates that it built. It does have a genuinely working deliberately narrow no check transpiler from a five file corpus, but everything else is code that hasn't been proven and the tests are currently failing. Currently, I have a bunch of work that was stuck in a stash, so I'm going to clean that up and figure out how far it actually is of the Obviously, I haven't had time to actually play with this, so we'll see where that ends up.
I'll show more in a minute, but I want to talk about the other things I worked on first. If you saw my video where I dropped all the ideas I wish someone would build, you might remember the idea of a Dropbox-like cloud that lets you sync your dev folder across machines. You probably understand why I wanted this now cuz I had all this work across all these computers, but I also decided to just throw out a goal and see what happens.
Before this, I hadn't actually analyzed how many tokens that one goal run burned. Turned out it was 71.2 billion, which at the fast mode pricing would have been around $91,000. I had not went to go check and see how well this works. I haven't even opened the project. I just let the goal run nearly forever, and I guess it did. I actually learned that it was using PlanetScale when one of my friends who worked there mentioned, "Yo, I heard you were working on FS2 again.
We saw you registered on PlanetScale. How has it been?" I was like, "Wait, I did?" And then I went and checked the logs and I saw it had autonomously registered for PlanetScale itself in order to build what it was trying to build. I'm actually much more excited to go try this out now cuz I did not think it would get as far as it did. So, uh fingers crossed it went well. I'll be sure to check it in the near future. I have to get it actually deployed though, which will be a lot easier because the computer use is phenomenal.
So, I'm really, really excited to see if it can not just set up this project for me, but actually go into all the dashboards and set up all of the things it needs to really deploy it. Ooh, the bootable Codex drive was actually one of my favorite things I did. Since I was managing all these machines on my network, I needed to make sure I could actually set them up properly. And I was getting tired of reconfiguring whatever Linux distro over and over again, and also having to recover it when they broke.
When I was ripping an SSD out of one of the machines, it didn't properly reconfigure the boot partitions on the existing SSD that I left in the machine. All of the content it needed was there, but it wasn't actually pointing at the right partitions, so it open up an empty grub that could only run memory tests. I realized I wanted an easy bootable drive that I could use with Codex and Claude already off, so I could plug in, boot to that, and then use it to fix things.
But setting up that image was taking a while. So, while I had one instance of Codex configuring this flash drive for what I wanted here, which it did, and it's awesome by the way. I highly recommend setting up a drive like this, where you can boot any machine to it. Not a Mac, of course, cuz Macs don't boot to Linux nowadays particularly well. But all the other machines on my network, I could plug in the drive, reboot, boot to the USB, connect the network, and now I have remote access to that machine with Codex and Claude already set up.
Super cool. But I didn't even end up needing it for the box I was trying to repair, because once I booted it, I was asking Codex for help like getting around weird BIOS things cuz HP's BIOS is trash. It kept hallucinating things I could do in the BIOS. So, eventually I said, "Okay, here, have direct access." I used my remote KVM, which lets me connect to the HDMI and a fake keyboard and mouse, so I can access the computer fully through the web, including the actual BIOS and config, and all of those types of things.
It lets me reboot a machine remotely, super helpful. But, you also have to time a bunch of presses in order to do things like get to a specific boot menu or get around grub. So, I just figured I would have to do those parts, and maybe Codex would be able to help. When 5 6 started hallucinating those weird BIOS things that didn't actually exist, I told it to use the tab in helium and look and see and click around to see if it could find what it thought existed.
And it failed. And then it rebooted, and it got into the broken grub state. So, I told it that it was in that broken grub state and asked for help fixing it. And it effectively said, "Hold my beer." Rebooted a few times, got into a shell in grub, booted correctly, and then used the computer used to remote control the computer, opened the terminal, and then fixed the boot partitions. All autonomously. These are the types of tasks that used to give me nightmares.
I know how to do them, mostly, but they require a lot of research and patience and risk tolerance. 5 6 just went out and did it with no additional effort. I still struggle to believe that it was actually capable of doing that. It was one of the most like, "Oh, we're really there." moments I've ever had, and it made me start pushing computer use way harder. And it did all of this faster than it was able to set up that recovery boot drive.
Both of which were unbelievably cool, by the way. I use both a lot. I have computer use through the GL.iNet KVM, which is still one of my favorite devices. I'll have a link in the description if you want to grab one cuz it's so helpful. But, I also got the boot drive, and I can use both of these to configure my network and all the machines on it. I also have a repo that I made called fleet that organizes all of the different computers I have, what their purposes are, what's configured on them, and what my default configs are.
So, when I get a new machine, I just open up that repo on this one, I tell it what I got and why I want to set it up, and what I want it to do, and then it uses all of that context to configure that machine exactly how I want without me even having to touch it. It has been so nice. I keep remembering random things I had the model do that impressed me. Like, I had it overhaul skate bench again and it did a great job. It's way more reliable and easy for me to control now.
There is still a bug with the reporting for token spend here cuz I know that those numbers are very very wrong, especially when you compare it with the costs. Not very good skate bench score to be fair, but that's not what we're here to talk about. I'm here to talk about how useful these are for coding. I had it build a 3D version of fish slop that I currently can't find on my machine, so I'll have to figure out where that is later.
I am sorry. I was really thinking I'd be able to grab that, but apparently when I CD'n it just doesn't exist. Oh, I had to help a bunch with prime day stuff. I do a prime day thread cuz I am a nerd for hunting deals and it was able to verify a lot of the deals I found, but more importantly create genius links, which is my link tracker that makes it easier to redirect to the country of choice for whatever country you guys are in.
That's normally a very slow manual process clicking around the web. It automated the whole thing through browser use for me and responded with all of my links ready to go. It was so nice. It also wrote some copy for me that was absolute garbage, so I threw all of that out, but the rest was impressive. It built my little terminal token tracker, which I'm sure Phase can sneak a picture of somewhere in here. I showed it off in the podcast, which by the way, if you want to hear me yap about this model even more, I have a whole podcast episode where I do that, too.
Filmed with Ben, who also had early access weeks before the model was officially announced. Very fun episode. I have learned a lot more about it since, but I think that was a good like oh podcast episode for us where you get to see our genuine reactions. But I set up that terminal with all of the content on it and its ability to get that content from my fleet using this model. It did a very good job, but its UI was incredibly ugly and I had to rebuild what it did there with Opus 48.
Spoiler, this model isn't magically perfect at front end like other models can be. It's much more steerable and has less bad of taste by default, still not a great front end model by any stretch of the imagination. I was wrong. Turns out I do actually have a version of to slop here that I haven't seen before. It generated all of the art itself, too. It's not great, but it's passable. UI is a bit garbage though, if I'm being honest.
Want to find the 3D version though. Give me a sec to do that. And here is the 3D version where it made everything itself. It made the 3D environment. It made all of the rocks on the ground. It made its own textures. Some of the textures are better than what I got out of something like Grok, but not all of them. I think the rocks are actually quite a bit worse. It does get the controls a lot better though. I had to push it hard to make the controls more like a Subnautica style like adventure where you're steering and driving your ship.
It also screwed up the shoot controls a bunch and it made some of the weirdest most hideous models for the monsters and for the fish. It It's understanding of 3D is uh the best I can put it is strange. But uh yeah, this is one of those areas where like the progress is real. It's still not at the point where I would want it to ship this at all, but it is progressing fast enough that I'm definitely keeping my eye on these types of like 3D game use cases.
I did a bunch of other weird things like trying to create my own Presto alternative that was more minimal. Ended up deciding on just using pure instead directly. I had it create ClaudeX which let me use my Codex off in Claude code so I could test using Claude code's workflows with Codex. And it was really cool. Definitely a thing I'll talk about in the near future when I start comparing the models more directly. Ooh, this is one of my favorites.
I created a little app that is hosted on a Mac mini at my office that checks which Mac addresses are on the network to have a little tracker that shows who's currently home or at the office to see who's in. Makes it really easy to know like what to expect when I walk into the office. Oh, and I had it overhaul my plan service that I use for hosting all my HTML plans including the one we're looking at now. It ended up building a bunch of really useful stuff for it.
I'll be honest though, I had Fable do the majority of the recent cleanup work cuz I only had access to Fable for a bit. But, yeah. I know that was a lot. I'm looking at how long I've been recording for and I'm a little scared, but I wanted to show what I've been doing with the model so that in the future when I talk about what the model is and how it works and what impressed and didn't impress me, you have the context of the things I actually did with it.
Spoiler though, those future videos are going to be even more interesting. This model got me genuinely excited to build more and think bigger. It's without question one of the bigger reasons I gave the talk I gave recently at both Cascadia JS and AI Engineer about thinking bigger because this model pushed me to my limit. I had to push it to do bigger, harder stuff because the simple things were too simple for it. I don't have to steer it the way I did with 55, I just tell it go and it goes and goes and goes without needing help managing context or getting not getting lost.
It's just kind of capable. It's a workhorse and I made it work hard. That said, I'm going into this without any idea how others feel and I'm curious how y'all do. Let me know in the comments and keep your eye out for the follow-up videos and again, make sure you're subbed and you hit that bell because there's going to be a lot of content about this model coming very, very soon. Thanks for sticking through to the end and until next time, peace, nerds.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.