Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Theo - t3․gg · @t3dotgg
Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
10:086.6x the video's typical replay level
after my rant about the current state of Codeex. Oh boy. So as I said the sub aents version being used doesn't expose the necessary functionality. What do I mean by that? Well what I mean is that there are two
Said at 10:02
Most replayed moment #2
3:386.4x the video's typical replay level
for no money. And if you do need some money, they can help with that, too, because they're currently hiring. So, if this is interesting, hit them up. Give the whole web tier your agents at soy.firecrawl. First and foremost, let's start with what is ultra. I'll start in the terminal so I can show you when you select
Said at 3:31
Most replayed moment #3
25:285.2x the video's typical replay level
making this awesome demo of what the model selector should look like in Codeex. Sorry, chat GPT. You have the different models in different effort levels, but you also have ultra as a switch on the side. Because again, ultra isn't an effort level. Ultra is a skill.
Said at 25:20
The graph counts replays. It does not show where viewers stopped watching.
Words
5,276
Runtime
25:58
Speaking pace
203wpm
Reading time
22min
203 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
When OpenAI put out GBD 5.6, they didn't just put out a new model. They put out three. But they didn't just put out three new models either. They also put out two new reasoning levels. They added max and ultra as options you can use in tools like Codeex. And I'm here to tell you something about those options. As good as OpenAI makes them sound here, saying that it's the highest capability setting, coordinating multiple agents across parallel work streams, I'm here to tell you that they're kind of lying. Not because Ultra is bad, because Ultra isn't a reasoning
102 words, the words spoken in the first 30 seconds at 203 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 334 |
| Average words per sentence | 15.8 |
| Longest sentence | 107 words |
| Questions asked | 14 |
| Sentences containing a number | 43 |
Most used terms
Filler phrases
44 in total: like 25 · actually 7 · I mean 4 · you know 3 · kind of 2 · basically 1 · literally 1 · right? 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
When OpenAI put out GBD 5.6, they didn't just put out a new model. They put out three. But they didn't just put out three new models either. They also put out two new reasoning levels. They added max and ultra as options you can use in tools like Codeex. And I'm here to tell you something about those options. As good as OpenAI makes them sound here, saying that it's the highest capability setting, coordinating multiple agents across parallel work streams, I'm here to tell you that they're kind of lying.
Not because Ultra is bad, because Ultra isn't a reasoning level. It is also bad, which we'll talk about plenty, don't worry. But the thing I want to make sure is clear is that Ultra is probably not what you think it is. But also, it is bad. I am incredibly frustrated with the way Altra has been rolled out. Both because it was not available to me during my early access testing, so I had to build all my familiarity with it in the last few days, and also because it's caused a ton of confusion to people who are trying to use these models the best they possibly can.
And ready for the most unexpected twist here? I blame Anthropic for all of this. OpenAI as well for shipping something that doesn't work. But Anthropic are the ones who started this trend that is not very good and is misleading people like you and I into using things incorrectly and not really knowing how to talk about them. I have a bunch of questions I want to answer in this video. What is Ultra? Why is it bad? How can it be fixed?
And most importantly, what the is going on with codecs? because I am at the point where I am now using 56 soul in other harnesses. Specifically, here I am using it in claude code. I want to do my best to break down what's going on and why I'm making all these changes after a real quick break for today's sponsor. There's a weird contradiction I've been noticing with AI. On one hand, it benefits greatly from having access to things on the internet.
But on the other hand, thanks to AI and all of the scraping that's been going on, most of the websites and sources that I need to get that info from have locked down and are nearly impossible to access. If only there was a really good open-source solution that would allow your agents to access data in formats they understand from all over the web. Oh, is it on the screen? Yeah, wire crawl is that. They built all the open source tooling your agents need to get data out of the web.
And they also offer a super generous hosting platform for doing it yourself. I know most of us aren't reading outputs from APIs anymore, but they really are this simple. You put in a URL, and what you get back is a markdown version of the page, a JSON version of the page, as well as a screenshot of the pages content. They have SDKs for pretty much everything you're already using. And if not, you can always curl or use their CLI, which honestly, the CLI has really, really impressed me.
Once you paste an API key, you can just tell your agents about it, and it'll give them access to info they might not have otherwise had access to. And the result is super simple markdown that your agents can actually understand and use for things. If all they offered was this direct URL scraping, it would already be one of the most useful things in my toolbox. But they offer way more. They have general search for finding things on the web that you might not already have a URL for, as well as interaction for when you want to actually do things on a page.
My favorite though is the monitoring feature. Monitor lets you get notified when something changes on a page, like maybe a price goes down or a release happens or an early access thing comes through. It's so useful for you and your agents to know when things change on the web. And as great as the API is, not every tool can use it, but you know what? They can use an MCP server. Do you know how useful this is for Chat GPT and Codeex as well as every other tool ever.
It's incredible. Two more really quick things they wanted me to make sure you know. The first is that you can use this without even having an API key. It's just free. You'll hit limits eventually, at which point you can go add an API key, but like it's insane how generous all of these tiers are. You get a ton for no money. And if you do need some money, they can help with that, too, because they're currently hiring. So, if this is interesting, hit them up.
Give the whole web tier your agents at soy.firecrawl. First and foremost, let's start with what is ultra. I'll start in the terminal so I can show you when you select the model that Altra is presented here as a reasoning level. You have low, medium, high, extra high, max, and ultra. If you use the chat GBT app, formerly known as the codeex app, they did an interesting thing here where it starts at the bottom as terraite, then it goes to soul light, then soul medium, soul high, soul extra high.
This is a change. It was going to ultra before and I think they actually listened to me and they hid ultra from that slider. I shouldn't have installed the update I just installed because before I installed the update, ultra was at the end of this list. I have told OpenAI verbatim multiple times now that Ultra should never have been included the way it was in this selector. It even had a fancy purple gradient it would do when you did it before.
I'm thankful they killed that cuz that was the biggest mistake. Why is it a mistake? It's cuz Ultra isn't a reasoning level. I'll explain the easiest way I know how to, which is with Claude code. Not because I'm going to ask it to explain, but I want to show you what happens when you do slasheffort. You have options here. Low, medium, high, XH high, max, and ultra code. Huh? Ultra code. Ultra Ultra Code. I wonder if there's some overlap here.
Ultra code was a pretty cool feature in Claude Code, but you'll notice underneath here, Ultraode says XHigh plus workflows. Hm. If you haven't been keeping up with my recent procla code arc, this is the big part of why it's workflows. Ultra code isn't workflows, but it's a way to trigger them. The more important piece to look at here is the other thing it says underneath, which is X high. Your options are high, XH high, max with the fancy little like gradient on the text there, but then ultra code does the super fancy gradient.
Silly, but watch at the top here. I currently have Fable 5 with high effort selected. So, when I press enter here, it's going to change that to ultra, right? Nope. It changes it to Fable 5 with X higheffort. And ultra code is just a state in the bottom. The reason for that is that ultra code isn't a reasoning level. Ultra code is a skill and ultra is two. What do I mean by that? What I mean is that it's the same as when you do a slash command to pull in a skill where it takes the content of that markdown file and it appends it to the system prompt so it's there when the next message comes through.
The reason they do this is they want to make it easier to trigger sub agents. Ultra and Ultra Code are both instructions for the model to use way more sub aents. In Cloud Code, this works through workflows and in Codeex, it works through the agent implementations called V1 and V2, which I will crash out about momentarily. Don't worry. I just wanted to make sure that we understood this first and foremost. Ultra isn't a reasoning level.
It's effectively a toggle that turns on a change in your system prompt, telling the model to do more sub agents. As much as I love to give anthropic they did mostly get this right. It's a little too easy to leave Ultra Code on, which isn't great. But as an implementation with the system prompt change and specifically the default to X higheffort, not to max, when you go to ultra, the reasoning level is lower than it was on max.
Again, when I go to slasheffort, low, medium, high, XH high, max, ultra. And the thing that's unintuitive is that Ultra Code actually points at X high, not at max. This is not what Codeex does. Sorry, this is not what Chat GBT does. I also can't help but notice they are hiding the max option now. They're listening to me. It's fun seeing in real time OpenAI making changes based on the things that I'm shouting about on the internet.
But Ultra is still hidden here under effort where it never should have been in the first place. When we click it, you'll see all of these other changes under the hood. Ultra is using max reasoning levels, which is absurd because max reasoning burns tokens. I talked about this a bit in my previous cost savings video, but the TLDDR is that max does up to two times more token burn than X higher high. Depending on the bench, it's between 4% and 10% better at more than 2x the cost.
So, I don't think max makes sense almost ever. But ultra isn't max. Ultra is nearly infinitely recursive maxes, which means it burns tokens. This is the primary reason Ultra is bad. The problem is that the parent agent and all of the sub aents are set to max reasoning. Well, Theo, this is an easy fix. Just tell the parent agent, the top level, to spin up the sub aents on a lower reasoning level. I'll have more to say on this in a bit, but the TLDDR for now is that the version of sub aents that soul uses does not currently allow for selection of reasoning levels.
So, if you have ultra set, all of the children sub agents spawned are also going to be ultra. I crashed out pretty hard about this on Twitter going as far as to say that claude code is far ahead. At the very least, they should give me an interim solution where I can hard set sub aent reasoning levels. The current implementation is so bad that I think it has caused this perception shift where people think that the new models are way less efficient and will destroy your rate limits.
The new models won't do that. At least with a little bit of careful use, they won't do that. Ultra absolutely will. I blew my first 5-hour limit in 20 minutes using ultra on fast mode cuz I didn't know how much crazier it would be. Instantaneously evaporated my limit. I burned a manual reset and I evaporated again another 40 minutes. I hit the 5-hour limit twice in under an hour. All because I had one ultra run going.
Hitting your 5-hour limit in 20 minutes is bad, but what would be much worse is hitting your weekly limit in an hour and a half, which is dangerously possible right now because OpenAI is temporarily removing the 5-hour limit due to people hitting it too fast. This is a good change overall, but if you're using Ultra, it's a little scary because previously that 5-hour limit being hit is almost a warning, like a hey, be careful. you're going to overuse this and lose your weekly limit.
And since the five hour is like 20 to 25% of your weekly, you would only lose 20 to 25% of your weekly usage. With this change, it is very much possible that Ultra could nuke your whole week. So, be careful. I'm realizing now that the order of events I planned here isn't actually going to make a lot of sense. So, I'm going to move how it can be fixed to be after my rant about the current state of Codeex. Oh boy. So as I said the sub aents version being used doesn't expose the necessary functionality.
What do I mean by that? Well what I mean is that there are two versions of sub aents in codeex right now labeled v1 and v2 internally. The codec cli which is what powers the codeex desktop app is open source. So we can actually look through the code here. I have 56 soul in codec and claude code breaking down how the sub aent implementations work. so that we can compare based on their findings. Codex finished first. I'll look at cloud codes for the difference in a minute.
I had them both going through the actual code base to try and break up the difference between these implementations. Obviously LLM speak so we'll get through and I will do my best to clean it up. V1 is like a dispatcher hiring temporary helpers by ticket number. V2 is like a named project team with an org chart and mailboxes. Good enough. V1 is very much a simple tool call that is a way the model can trigger another submodel to go do something.
So let's say you tell the model to go review three PRs. It can spin up three sub aents, one for each PR with the instructions to go do that one thing and then when it's they're done, they'll all send up their findings and that top level agent will decide what to do from there. Side note, this page is ugly as sin. I have things to say about that soon, don't worry. But it starts here with five definitions worth knowing.
They refer to the parent agent as the root agent. I think that's fine. Then there is the sub agents which are the things spawned by that root agent. Sub aents can also spawn themselves. We'll have to deal with that in a bit. Then there is the context which is all of the info that the top level agent and the sub aents have. Usually sub agents are spawned by writing a new prompt or summarizing the stuff going on in the main thread so that the sub aent has a limited set of context.
Remember that because there are things going on here. The mailbox is a new concept in the V2 implementation where there are messages for the team as well as the ability to send messages between different agents. And then there is slots which is the number of agents that can be running and the shared workspace which is the files the agents are working on. This does not need to be in here. This is not a great description.
What a surprise. The very least it made decent diagrams here with V1. you have that root level agent and it can spawn sub aents to do specific things that return a result when they're done. V2 is a complete overhaul where the different sub aents can talk to each other and spawn sub aents of their own. This sound familiar? Remember that infinite recursion thing I was talking about with Ultra? Yeah. There's a problem though.
V1 is the finished implementation. This is what Codeex uses by default. And V2 is an overhaul that they are still working on. It's still a work in progress. V2 is very much unfinished and if you turn it on, you get errors for having V1 at all. Especially if you have anything changed about V1 like custom limits, custom instructions, whatnot. You can't have V1 and V2 on at the same time until now because it's a really really annoying exception.
In the models cache JSON file, which is the file that Codeex uses to determine which models are available to you, there's a new field they added multi- aent version, which is V2 for Soul and Terra and V1 for everything else. So now it doesn't matter what you have configured or if you've manually opted into V2, the new models are always routing to V2. But V2 has some other quirks. With V1, Codex had access to this handful of tools for spawning, sending, waiting, closing, and resuming sub agents.
They would create a separate thread effectively with its own context for the agent to go do its thing and then respond when it's done. The parent would wait for the agents to be complete and then summarize whatever it got back. V2 is quite different because now it is breaking down tasks instead of just sub agents. An agent requires a task name. A root child named research becomes /root/ressearch. its tests become slrootressearchests.
So this is a pathbased naming scheme for agents to spawn sub aent layers. But the biggest issue for me is the way it handles context sharing. By default, all of the history in your main thread is going to be shared to all of the sub aents. Spoiler, this is really stupid. While 56 is way better at not falling for context pollution where some bad thing in the history affects how it behaves, it still does. This also is a massive increase in cost because the sub aents now have way more context.
I would assume that part of why they did this is to keep the cache consistent between the parent agent and the sub aents, but I'm also pretty sure there are changes in the system prompt which is slightly higher, therefore busting the cache. Not positive about that, but I would be surprised. Either way, getting one cache right at the start of a new sub agent thread is far from a high cost. I think it's stupid to try and preserve cash between the parent or sorry, root agent and the sub agents.
So, yeah, dumb. I don't like this at all. As I said, by default, it now takes all of your turns in all of the history when it goes to that sub agent. But you can set different amounts to share, like none or a number like three. And as I said, the default is the full history, which I think is really dumb. It also filters out the tool calls. So that breaks cash. Where things start to get really weird is this idea of mailboxes.
Instead of just having a sub agent go do some work and then respond when it's done, they now can send typed messages between each other. Send message cues a note. Follow-up task gives an idle helper more work and can start its next turn. This allows for the parent agent to send additional work down to an existing sub agent when it decides it wants more done. This is one of those things where it like does genuinely sound really really cool.
In reality, mixed results. The finished work is routed to the direct parent. So again, if they're nested, tests sends its results to research, which sends its results to root. Yeah. And waiting no longer means wait till it's done. It means wait till we get a message back. The root agent can also list all of the sub aents and even interrupt them if for some reason it wants to. If it got info from another sub agent, it's like, "Oh, that thing probably doesn't matter anymore.
I can kill that." There is no depth limit in V2, so it could spawn infinitely nested sub aents, but it can only run four at a time by default, which I had changed, which is a big part of how I hit crazy usage so quickly. Random observation, I did run this prompt on both Codeex and Claude Code. Claude Code took a bit longer. Again, I was using 56 soul on both. If you feel like this design's a little bloated, you're not the only one.
I love this message from JKF here. It sounds like too many people designed sub agents v2 and caused it to be bloated. I agree. I don't really like the design of V2, and I didn't use it much when I was testing 56 soul because again, they hadn't added this change where it would default route to V2. I did turn it on manually near the end of my testing out of curiosity and I noticed it would work a lot longer. Didn't get far enough on real work to really see a difference in the quality of the output, but it definitely goes a lot longer and spins up a lot more.
Don't necessarily love that, but it can. All of that said, I've been using it more since and I'm not really impressed with the quality of the work and output I get from V2 with the current settings and setup in codecs. While I do love the ability to spawn sub agents to break up work and control the context a bit more, the changes to how context is managed and the changes to how messages are passed makes this noisier is the simplest I can put it.
And it also burns way, way, way more tokens as a result. And this leads us to our final question of how can it be fixed. I know this stuff scares a lot of people like they're afraid if they don't get it right initially, they might miss out on this opportunity. they might fall behind all these things and a lot of you guys are trying so hard to stay on top of the best solutions. If you're here out of fear and not out of excitement and curiosity, the only thing I want you to take from this is that you should just wait a bit.
The current implementation is bad and has rough edges, but the OpenAI team is incredible at taking feedback. I have been shocked as I film at just how many of the things I'm complaining to them about have been addressed between the sixish hours ago where I sent the complaint on Slack and now when it's not there anymore. If you just wait a little bit, these things will improve. It continuing to use the defaults and trusting that these companies will get it figured out eventually, you'll be fine.
It's not that big a deal. I know I am crashing out hard about this, but I want you to realize this is because I am a hardcore nerd who cares too much and is just digging into the details because they want to. You don't have to do this to stay ahead. You can just use the tools as they work by default and they are still really good. With all of that said, I think OpenAI needs to rethink how they do sub agents. As I mentioned before, ultra in codeex is a pretty blatant copy of ultra code in claude code, at least in how it is meant to work where it's added to the reasoning effort slider.
It is meant to trigger sub aents and be a way to force sub aents on. It basically just appends to the system prompt. Please, please, please use sub aents. But in Ultra on codeex, it is doing that through these weird sub aent tools like we've seen many times before. In cloud code, it's using workflows. I've talked about workflows a bunch before in my Claude code is good video and I plan to talk about them even more in my Claudeex video where I show how I'm using 56 soul in cloud code directly, not as a sub aent or a thing that cloud code calls literally as the model inside of cloud code directly.
It's way better than I thought it would be. The reason I'm doing this is workflows. To put it simply, workflows are a way to programmatically define a bunch of sub aent stuff. In cloud code, you can actually save the file it writes because it's making a JavaScript file to do this work. And that's by defining a meta, which is a name, a description, and different phases because it has stages throughout its work. It then has schemas, which are the typed outputs that each of the phases should have.
It then writes a prompt, and this is programmatic. So if it wants to insert things programmatically, use a map or some type of loop to add context in and do different steps, it can because it is just code. Once this is all defined, it can call the different phases. Sorry to my editor phase as I'm talking about phases. I'm sure that won't be confusing at all. So it starts phase reviewing. It defines reviews as a bunch of parallel agents.
We have the first one here, which is a hard-coded prompt. Your assigned perspective is 56 soul at high reasoning effort label review 56 soul phase reviewing model 56 soul effort high schema review schema agent type general purpose and we do the same here with another one that's using terra as well and then one here using fable cuz I asked it to use all three of these once this phase is completed because remember we awaited it up here I know crazy we're reading code in a theo video again it starts the synthesizing phase where it awaits another agent call where it is passed these review results s as a JSON blob that's been stringified.
And now another agent using Fable 5 on high is going to process all of the data it was given and respond following the synthesis schema. And this is how you can programmatically use sub aents. All of the stages are defined ahead of time. And depending on what the outputs of a given stage are or a given agent run are, it can decide if it should put the results in another run in another sub aent in a different stage. This one is really, really simple because there was only two phases.
There's the review phase and the synthesizing phase. I've had many more complex ones, though. I've had some that had 12 plus phases. Here's one that had five. It was research, verify, synthesize, critique, and finalize. Each of those got a schema for how the output should be formatted. It was given a common prompt to append to the top of all of them because all of them need to know what the directory is and a bit of other information.
It created these different topics that are keys and prompts for different things that we'll be doing throughout. It then starts the research phase. It goes through and defines multiple sub aents to do all of that. It clears out any that didn't have an actual response in case that happens. But you can also filter based on a given field. Like if you had a field in the response that was should keep researching true or false or needs another review or needs followup or solves problem whatever it decides to put in here.
It can now use that content to determine where to pass off the work next because it is code. The problem with most sub aent implementations is that they're leaning too hard into tool calls. They work because the model will spit out a call spin-up sub agent with this information and then it does it and then it does another tool call for another and then it does it and then it does it again where the agent has to keep all of this context itself and spin everything up as part of the LLM's work.
With workflows it's hardcoded but it's hardcoded on the fly. Previously, sub aents were just full-on hard-coded where certain harnesses cough cough, oh Mike, pi, cough, cough, open code had hardcoded sub aent descriptions like this is a researcher sub aent. It does these things. This is an implement sub aent. It does these things. Workflows is a really cool in between where on the fly it will create these different sub aent types and classes and what they're expected to return with and then create a flow, a workflow if you will, where each stage is determined based on what happened on the fly with code that was written ahead of time.
This does a few things really well, but the biggest it does is it kind of puts a hard cap on when this all ends. Instead of just going forever, which is really possible with Codeex's Ultra implementation, there's a fixed number of phases. So eventually it will get to the end. It might spin up a shitload of sub agents in one field. For example, if you ask it to read three files and find things that are worth fixing, it might find 72 things to fix in the first file and it might spin up 72 sub aents in the next section, the next phase where it is fixing the things that it finds.
That's fine though because it will eventually get to the end always. Ultra is much more likely to just go forever because it ends when the agents decide to end, not when the code is completed. So, what I'm trying to say is I think OpenAI copied the wrong parts. They copied the UX, which was bad, where they disguised it as a reasoning level where it isn't, and they didn't copy the implementation, which is good, because workflows are an awesome way to break up work programmatically.
They are so awesome that I'm causing problems for a lot of people because of it. And I have a whole dedicated video coming up about my usage of 56 soul in Claude Code. It's probably going to be the next video I post. So, make sure you subscribe and hit that bell if you want to see it as soon as it drops. I think you'll be surprised. Spoiler, I'm going to spend the majority of this video crashing out about the system prompt in Codeex because it is so much worse than I thought it would be.
One last thing, shout out to Maria for making this awesome demo of what the model selector should look like in Codeex. Sorry, chat GPT. You have the different models in different effort levels, but you also have ultra as a switch on the side. Because again, ultra isn't an effort level. Ultra is a skill. And it seems like everybody from TBO to Dominic to even GDB himself agrees that as silly as this is, something like this does make sense.
So for now, don't use ultra. If you really want to get this type of thing, keep an eye out for my video on how to use workflows in cloud code with GPD56 soul because even Tibo supports my chaos here. Hope this [snorts] was helpful and until next time, peace nerds.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.