Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
Most replayed moment at 13:49
2.3x that video's typical replay level
Fable uh and it runs into an unknown, ask it to log it, right? So that um you uh you can see where the deviations happened and then you can sort of figure out why as well, you know? It will usually give you some context about what happened.
Said at 13:43
The graph counts replays. It does not show where viewers stopped watching.
Words
9,944
Runtime
1:01:58
Speaking pace
160wpm
Reading time
41min
160 words per minute, between the 160 25th percentile and the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Hi everyone. Thank you for being here. So, today we're going to talk about Codex. My name is Katya. Katya Gribuzina, and I'm with VB. Uh we are both uh working in the developer experience team at VB. based in London, and so our role is really to help developers build and get the most out of our products, uh including Codex. And so, today we're going to start with a quick Codex overview, uh just so we know how
80 words, the words spoken in the first 30 seconds at 160 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 520 |
| Average words per sentence | 19.1 |
| Longest sentence | 129 words |
| Questions asked | 33 |
| Sentences containing a number | 32 |
Most used terms
Filler phrases
728 in total: um 232 · like 154 · uh 152 · you know 95 · sort of 49 · actually 18 · right? 15 · kind of 5 · basically 4 · literally 4.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
Hi everyone. Thank you for being here. So, today we're going to talk about Codex. My name is Katya. Katya Gribuzina, and I'm with VB. Uh we are both uh working in the developer experience team at VB. based in London, and so our role is really to help developers build and get the most out of our products, uh including Codex. And so, today we're going to start with a quick Codex overview, uh just so we know how many of you here are using Codex.
Can you raise your hands? Yay! Okay, cool. So, we we're not going to stay too long on that overview part, uh and then we're going to we're going to do some demos. So, we're going to show you plugins and automations. Uh VB's going to talk about sub-agents, and then about the bleeding edge. So, hopefully for those who already know Codex and use it, uh you'll learn something that uh you didn't know about. And then we have some time at the end for Q&A.
So, um feel free to ask anything. Also, this is a a workshop format. So, you know, if you have like a a pressing question, um feel free to ask. And I see you all have your laptop, so also feel free to kind of follow along with us. We're going to show you how to do some things, and you can like try it at the same time. And during the Q&A, uh that's also like the perfect time as well to try things on on your side. Okay.
So, um to start, just for those who don't know Codex, or even if you know it, maybe you don't know it that well, uh Codex is our open AI's is open-eyes software engineering agent. So, it's not just a coding agent. It's not just a, you know, uh, an agent that writes code. It can do much more than that. It can run commands. It can run tests. It can explore code bases. It can really do everything that a software engineer would do.
And so, it's based on our models as a foundation. So, for example, GPT-5.3, Codex was, uh, our previous ones. We also have the Spark version, which is like the super fast, uh, model that that we have. The the state-of-the-art model right now is GPT-5.4. And we also have a mini version that, uh, came out last week. And, you know, every time we make improvements, every time we have better models, Codex benefits from it.
But, it's not just the models. On top of that foundation, we have what we call a unified agent harness that will manage, uh, evaluates the agent's behavior. And that is a wrapper for tool execution, for environment setup, for everything that, uh, can let the agent, uh, do its work and run smoothly. There's also safety, uh, safety embedded in that harness. So, all of that is Codex. And then you can interact with it through different surfaces.
So, you have the Codex app that we're going to talk about in, uh, in a few minutes. Uh, you can also interact with it through your IDEs with the extension. You can interact with it through the CLI, and also through other surfaces like Slack, for example. At OpenAI, we all the time just like ping Codex in Slack and ask it to fix things, or in GitHub as well. Uh, and on top of all of that, you can also integrate it with your preferred tool so that it can really, uh, work with everything that you're already using.
So, you know, you can integrate it with Figma, with Linear, with Notion. all of that combined uh, can let you really like do everything can let Codex do everything that a software engineer colleague would do. And so, as I mentioned, this is based on our models. And so, I'm going to let VB tell you a little bit more about that. All right. Um good morning, everyone. So, as we've been talking about the Codex app, the IDE extension, the CMI, and so on and so forth, all of these like harnesses as well as all of these services would not be nearly as good without the models powering them, right?
And just to sort of like take a step back. Back when I joined joined OpenAI, which is not really as far back along, was in December. Our leading model at that time was GPT-5.2. And and from there, we went on to release GPT-5.2 Codex, which was a specialized, you know, Codex variant of GPT-5.2, wherein we we sort of pushed how far you can you can take the model and you know, run it on long-running tasks, how far you can let it just continue chug along.
And then, shortly after, we followed up with GPT-5.3 Codex. Shortly after that, in partnership with Cerebras, we followed up with GPT-5.3 Codex Spark. And and most recently, we had we released GPT-5.4. And you can already see how we're sort of, you know, pushing this whole sort of model and harness flywheel as fast as possible, trying to bring the next next frontier as as fast as possible to you all, right? Something to note is and and and something that's not on the screen is at the same time we also, whilst we were pushing for larger models, which are really good for long-running tasks as well as really complex tasks and so on and so forth.
We also released GPT-54-mini and GPT-54-nano which you can use for short running tasks and sub agents which we'll talk about in a bit. And and something that we haven't really emphasized on this over here is is two things. One that as we as we sort of pushed on making these models better, we also worked quite a bit on making sure that these models can be served to you as fast as possible. What that means in principle is we we introduced something called web sockets which allow us to sort of create a connection between your your device as well as where the where the API resides to be able to give you roughly about 1.75x faster tokens without without really like paying the cost over to you.
At the same time we also released a fast mode which allows you to on top of the 1.75x get 2x more faster tokens. And and this is something which the team is continuously sort of hammering on. There's there's lots more speed improvements coming over there. And so to bring this all together at the start of this year we we brought together the Codex app. How many of you have used the Codex app? All right, that's a that's a fair good chunk of people.
To be honest, back in December and and and even before I was a hardcore CLI user. And um at some point during the app launch whilst I was beta testing it, dog fooding it, I you know, the Codex app became really like a really important part of my workflow. And the reason for that is it it brings together a really nice way to work across projects, number one. And number two, within within the same project work on multiple features at the at the same time.
The way you can do that is you can have individual individual projects like you can see on the on the left side you can have the Codex project, ChatGPT, Sora, and so on and so forth. But also within those you can you can use um work trees to work on individual sort of feature requests or you know, bug fixes or just like Q&A all at the same time without really interfering with individual tasks. This is something which we're quite proud of and you know, providing a native work tree support helps you do the same task and do multiple tasks at the same time without really having to context switch as much.
Um at the same time through the through the launch we've been trying to sort of increase the net benefit you can get out of the Codex app. And some of these features have been like you know, having a better automation support. Automation is is also something we're going to talk about just in a bit. But the short summary of automation is that you can essentially have a have a rough process that you want Codex to run. Let's say every day at 9:00 a.m. or let's say every evening.
Or let's say you want Codex to look through your calendar and and and create like a briefing for you. And that all is possible all within like the native Codex app setup with with automations. Um and then of course with the with the work trees and like more native get support you you can sort of work across projects and and just be able to you know, push changes as you want with whatever get persona you want to do it with.
Um last but not least, based on which surface you use you use the Codex app on. At the start of the year we released it just on Mac OS, but now we have uh native Windows support, which comes along with native Windows sandbox. Is there anyone here who's using Windows today? I'm cheering for you, man. And um so uh for the for the for the one gentleman over there, uh we have native sandbox support uh in Windows. We're the first of kind.
Um there is no other um you know, competing harness which supports like a native sandbox for uh Windows. Cool. And so, I've been talking on and on about like the Codex app. Uh been talking about, you know, all the models that we've been shipping. But what's what's new in terms of all the features um that we've shipped. Um This is in I think, if I'm not wrong, in the in the sort of descending order. So, most recently we launched plugins.
Um plugins is um is a way that you can bring together skills, uh MCPs, as well as prompts, and you know, any other thing really uh together in one bundle, and allow the model to do more nuanced uh matching while it's building. Um We also released recently mini models um which tie in quite well with sub-agents, which allow you to parallelize uh a particular feature or bug request or Q&A, whatever it may be, um at um at a faster rate all while it's making sure that you don't pay as much cost for your uh for your particular models.
Um and then we have like a bunch of other stuff which we're which we're going to talk about as we go through. Uh some of this is, you know, how you can uh how Codex is so good at like code review, how Codex is really good at security, um and so on. Um all of this um while we talk about all of this, um I want to sort of emphasize on this fact that um we're at OpenAI quite lucky uh that the community has really embraced Cortex.
In fact, just uh just last night we crossed the milestone of of crossing 3 million weekly active users. And this is a pretty big deal for us. And we we want to continue sort of supporting the developer community, the you know, enterprises, startups building on Cortex. So throughout the session, if you have any questions, please feel free to throw it at both Karthik and myself or even afterwards or just you know, ping us.
With this, I'll pass it over to Karthik. Thank you. And yeah, the the 3 million weekly active users thing is really it's really cool to see and it's crazy to think that it's also more than tripled since January. So just in a few months we've seen like huge adoption and yeah, and it's it's it's really really cool to see. Um Okay, so plugins. Plugins, I don't know if you've heard about it. It's it's quite new on Cortex.
The the native support for plugins. The idea of plugins, I'm going to show you what it looks like in practice and how you can use them, is that they bundle a bunch of things together. So like skills, apps, integrations, MCP servers and they they bundle that into reusable workflows. And so what skills, apps and MCP servers are, again, I'm going to show you, but just to introduce that a little bit. So skills are essentially reusable instructions packaged for specific processes.
So if you have something that you're doing quite a bit, you can actually create a skill for it so that Cortex knows about it. You can give it instructions, you can give it scripts as well, resources and all of that will save you from just repeating yourself over and over. So every time you have like a sort of neat workflow that is always the same, you can package that into a skill. You can actually ask Codex to create the skill for you as well.
And then apps are connections to other services. So, you know, again, we'll see a quick demo, but the tools that you use every day like Notion, linear, all of that, you can let Codex connect to it. And then MCP servers, you might be familiar with this already, but they basically expose tools for Codex to just extend its capabilities further, and it's tools from external systems. And so, all of these three things are already very useful on their own.
And what plugins do is that they bundle that so that you don't have to, you know, set up everything manually. You don't have to install multiple skills. You don't have to connect multiple apps. You don't have to connect multiple MCP servers. You can just add a plugin. And another thing I wanted to talk about in the Codex app, and that will show a quick demo for, is automations. Personally, this is like one of my favorite things to do with Codex because you can set up automations that run in the background, so like a cron job.
And you can connect apps, you know, you can use plugins there, too, and just set it to run on a scheduled time. So, for example, you know, you can set an automation for to run every day at a certain time, and it's just an instruction that Codex will run in the background. And last thing I wanted to show you with the demo right after is specific skills for web, app, and game development cuz we've, you know, we've heard a lot about developers who want to to use Codex to build these things, to build apps, to build games.
And every time, you know, they kind of repeat themselves, every time they kind of use the the same skills. So, we actually package that into specific plugins. And uh there's two skills that I want to highlight uh that are super useful and honestly that are game changer when you're developing something visual. It's um Playwright Interactive. And so, for those who don't know, Playwright is essentially like a a a headless browser, like a um you know, a sandbox browser uh that you can that Codex can just run and use that to see what it's doing.
So, you can open your app in a browser, and with the interactive version, you can actually click things and uh you know, just navigate your app. Um and and take screenshots and see the and analyze as they do screenshots. And then Imagen uh is a great way to just generate visual assets for your apps and games. So, enough talking. I'll show you a demo. Um I'm going to start by actually running this uh this one because this one is pretty long.
Uh when I ran it yesterday, it took like an hour to build. So, I also have like a the final version, but I wanted to show you like this this prompt, how Codex is going through it. And so, what I'm doing here is I'm using uh the Game Studio plugin, which is a again a bundle of a bunch of skills that are helpful for uh game development. And I'm asking it to use Imagen to create visual assets, so sprites for the games, and using uh Playwright Interactive to also debug the game and make sure that it works well.
So, we're going to let that run. And uh then we're going to talk about plugins for a little bit. So, let me switch to another project here. So, this developers website one. Okay, so this one is the repo for our developers website. Which is here. Sorry, I'm going to put that in full screen. And so, on our developers website, we have this page with all of the Codex meetups we have. So, you know, there's a lot. And all of that is actually in our repo, like in our code base in YAML files.
And so, I'm what I'm going to do is I actually added this Google Drive plugin here. You know, we have a a lot of featured plugin built by us that you can choose from. You can also, of course, add your own plugins. But I connected this Google Drive plugin that lets Codex access my Google Drive. And so, what I did is that I prepared this this spreadsheet called Codex Events with the event name, date, and city. And I'm going to ask Codex to just update this sheet with the current Codex meetups listed in the code base.
So, I'm going to start this. Again, it's going to take a little while. And so, let's check in on it. Okay, for the the game process still running. I'm going to show you when it's doing a little bit some more interesting things. But the last thing that I mentioned is automations. And so, automations is again something that you can just set up using apps. Do you can just ask Codex anything, but instead of it being interactive, like you're actually using the Codex app, you can set it up to run in the background.
So, for For some of the that I set up that are honestly helping me a ton in my day-to-day lives is one for Slack messages. So, I connected Codex to Slack and I'm asking, "Hey Codex, can you check every day at 9:00 a.m. the messages that I should reply to and flag if it's time sensitive or waiting for an urgent response? Can you also do a summary of all the things that have happened since yesterday on Slack?" And I'm asking that to bucket it to buckets per topic.
And then important information to be aware of. So, we have like important channels where company information, generally the things that you can that that that leaks in like one day, but so important company information is in there. And so, I just want to make sure that I don't miss anything here. So, that's the kind of stuff that I ask Codex to just summarize for me. Another one that is pretty cool is the is connecting Gmail and same thing like I receive honestly an ungodly amount of emails per day.
And so, I'm just asking Codex to check if there are emails that I should actually reply to and to check, you know, if it's time sensitive or if it looks legit or not because I do get a lot of requests that's I'm not necessarily something that I would uh that I would reply to, but this is like saving me hours per per day. And so, the way you can create automations is you can create it from here or you can also just, you know, uh say something like uh "Hey Codex, can you create an automation that will um look at Slack and look for anything that mentions Codex use cases and then list all of the important use cases that I should that I should put on our website.
So I'm going to let Codex think about this for a second. Should have used park. And it's going to come up with this you know, it's going to create the automation for me basically and I didn't specify when I wanted to run it but I can actually like Oh. Interesting. It's doing something different. Because this is a live demo so obviously it wouldn't uh Okay, normally it will it should like do a little pop-up so I can just like click on the Oh, it's doing it.
Perfect. It was just very chatty this morning. Okay. Interesting. Interesting. Okay. So Please create the automation. So this it should show a little pop-up if everything goes well but if not you can still like create it manually. Uh let's just see if it is doing it. Okay. I don't know what's going on but okay, let's just do it manually. So it will you can also create it from here and basically all you have to do is just call the plugins you want to you want to use you know, like use Slack and then choose you know the frequency where the automation should run, which project it should run in etc.
Okay, so let's check on our other tasks. Uh this one is still ready. Okay, it generated some some pretty cool sprites. We'll look at this after. Uh, let's check on our uh, task to update the spreadsheet. So, here CodeX took 2 minutes to actually analyze the code base. It found the source for all of the CodeX events where we have our YAML files and then uh, it wrote the 57 event rows. So, we have 57 events uh, currently listed on the website.
And uh, so let's check. Let's see our spreadsheets and yeah, we can see that it was updated. Nice. So, this is something, you know, this is a simple example, but every time you have something that's very, you know, uh, time-consuming and uh, anything that has anything to do with data, data review, for example, you can actually ask CodeX to do it for you. It has access to everything uh, on your code base and you can also feed it other inputs, you know, like other CSV files and then you can just ask CodeX to do that type of work for you.
Okay. Now, last thing, let's check on our uh, game. So, as you can see, CodeX is actually using ImageGen to generate I'm going to uh, de-zoom out a little bit. So, oh, nice. So, it's generating like all the sprites, all the game assets that I asked it to do. And this looks pretty nice. Uh, it's also So, it's going to take a while. Uh, what I'm going to do is I'm actually going to show you um, final results, but uh, as you can see like CodeX is just reading um, sorry, it's just generating all of these assets and then it's going to use the Playwright skill to see how that looks like in the app.
So, unfortunately, we don't have an hour to wait for the final So, let me just show you the one that it did yesterday. So, this is un uh untouched. Like, I haven't touched it. It's literally just Codex um who built this. And all of that like I had I gave zero input. I was just like do a platformer game with platforms made of bricks. That's it. And yeah, it generated everything. So, granted the the overall UI is not like, you know, I would probably iterate on that, but um I think the the platformer itself is pretty cool.
And what is really cool here is that literally like all the sprites like here, you know, I'm just like moving all around. And you know, that that's at least like five different sprites of the little character. And I didn't have to do any of that. You could also, you know, do a custom game with your face as input and have Image Gen just like create a a 2D version of you. Um so, that's a way that you can like leverage the Image Gen skill, the Playwright interactive skill, and that game studio uh plugin.
And just to show you what's inside like we have also the same thing for web apps, but it's a bundle of like all of these skills together. Um so, yeah, that's uh that's it for me. Uh I'm going to pass it back to VB. Thank you. Thank you, Katya. All right. Um Perfect. So, just to do like a very quick uh checkpoint uh and like a recap on what we've spoken so far. So, we went through like all the um all the models that power um the Codex ecosystem.
Then we went through all the surfaces you can consume Codex from. Um and then we went through uh uh plugins, how to use them, and what are some of the plugins that you can use. You can also create your own plugins um using plugin creator. Um you And And then we went through uh to speak about uh automations um and Imagen and and so on and so forth. Um Now, something to note is like as we as we continue sort of delegating more you know, more and more work on these agents.
It could be any of your favorite agents, uh Codex or not. Um one thing that um that you want to be sure of is whatever it is that your agent produces is of the utmost quality. Which means that um as we as we start sort of working on multiple features at the same time, multiple projects at the same time, it it's going to be quite likely that it's impossible for you to uh go and look through each and every line of code.
Which means that at least for the first pass, you want to have a way um which you can rely on um to review your code. And this is where um code review um sort of comes in. Um it's um by no means um am I bragging about this, but uh in my own biased way, uh Codex code review is one of the best in the industry right now. This is uh something which, you know, uh people on Twitter and LinkedIn um on our own uh sort of, you know, platforms, Discord, and so on and so forth, keep raving about uh that how is Codex code review so so good.
Um so, I wanted to spend like a quick hot minute on um on what it does. So, first of all, um it is available on the surfaces that you work at. Which means, number one, you are able to use Codex code review on GitHub. Um so, you can connect your ChatGPT account with GitHub, and for each and every pull request that you create, um you can set it up such that Codex can automatically review each and every pull request and it would typically give you you know some sort of a some sort of a you know uh What's this called?
A callout? Like this on the pull request itself saying that hey like this is something that is missing. Hey maybe you know P0 fix this P1 fix that P2 you know this is something that would be a good to have and so on and so forth. At the same time you can use a slash review on the on the Codex CLI or the Codex app and Codex will spin up you know large um sort of review process and so on and so forth. And very recently last week with my colleague Dom we shipped a cloud code plugin for Codex which allows you to you know essentially invoke Codex within your cloud code sessions to be able to get the same sort of state-of-the-art code review but in your cloud code sessions.
Um So um something to sort of see here is let's say that I am working on a project like this. By the way this is my this is my actual working setup at work. I this is like all which I work on. I'm not like everything that you see here is like all of these threads all of these projects is something which I work on day-to-day. So if you see something which you shouldn't just close your eyes. Uh And so typically what I would do is I would I would go through you know like a like a feature request or I would go through you know some sort of ask from from someone.
Um And um uh let's say over here I asked I asked Codex to do a bunch of things. So I'm just going to ask it to review its changes. Um And so then you get an option to you know either choose from a a branch if you have multiple branches in in the Git repo, you can choose it against a feature branch, against an eval branch, whatever it may be, and so on and so forth. Uh in this case, I'm just going to ask you to review um uncommitted changes.
Uh and what it does is if you see um here what it does is it spins off a totally new thread. Um and what that thread would do is um is it would essentially spin up a totally new Codex process which has uh like our own, you know, review system prompt. Um and it would continue sort of looking through not just the diff or like the list of all the changes, but it would also contextualize it with everything that is there on the uh on the model repo itself, right?
And so, a lot of the times um um Codex code review will like find find out changes which would have second order effects um which is not limited to just the, you know, diff or whatever changes you've made, but also to some other like modules which you haven't even touched in the pull request itself or in the changes itself. And this is um this is so effective that 100% of pull requests across all OpenAI repos um made by all employees um including Greg are are reviewed by Codex code review by default.
Um and that's when uh you know, that's the first pass that you take. Um Cool. And so, as you can see over here, um Codex worked for a minute and it came up with these with these sort of uh you know, updates like P1, you know, localize whatever revenue detail, P2, uh translate this to this, and um and so on and so forth. And what you can do like after this is um like essentially ask Codex to uh either like take a pass at fixing this or like open another sort of PR on the on whichever branch you're at, and then sort of go on from there.
Cool. Now, we get to sub agents, which is something which I'm personally quite excited about. So, first and foremost, what is sub agents? Sub agents is the is is essentially the ability wherein you can spin off a master task into decomposable, parallel, and independent tasks, which you can hand off to agents, which can which can allow these agents to sort of work independently, and then at the end of their run, get back to you and you know, give you a response.
And overhead like sky is literally the limit. Like you can spin up as many agents as you want. Of course, as long as your API key or your you know, whatever chat GPT Pro Plus gold subscription you're on can can can take. You can do a lot of like interesting things with sub agents. For example, what I'm doing on the screenshot on the left is I have a Codex agent repo, which we're going to look at in a sec. It's not public yet, but I hope that we'll be able to make it public very soon, which has a lot of personas for sub agents that you can use.
So, it's it's kind of meta. It's it's essentially sub agent personas like doc reviewers or you know, um test case creator or test case runner and and so on and so forth. And what I every now and then we would change the change the spec. And this is from before we wanted to change the spec of how how sub agents work. So, what I wanted it to do is to go through all of these 40-50 different sub agent personas, review them, and and and make sure that they're up to spec.
And of course, doing it without sub agents would have meant that Codex would open each and every file, and then review it, and then give me a summary, and continue doing it for like 50 different sub agents. In this case, um it it essentially created review slices, which means it created say, you know, these are the two uh files that um that, you know, sub agent uh Polly or sub agent Plato uh should, you know, uh essentially review.
And then they would spin up a new Codex environment, they would review those, and then at the end Codex will collate all of these, and um you know, give me back a response. So, let's let's give this a shot. Um, So, the repo in question is this. Um, it's um it's just a Codex Agents repo, which has a bunch of personas. Um, you can see that we have um we have quite a few sort of personas over here. Um, we've got like an accessibility reviewer, architect, and so on and so forth.
And this is like actually something which you can create yourself, and we're going to touch on that in just a uh in just a minute, is um you can you can define your own custom sub agents, right? Um, but think of this repo as like a collection of these sub agents. And um this is typically what you would have for for each and every sub agent. You would have a name, you would have a description, you would have a different sort of like, you know, sandbox mode, whether you want it to be right only, whether you want it to be read only.
Um, you and then you would have some sort of like, you know, instructions, um and so on. And so, now what I'm going to do is I'm going to ask Codex to um I'm going to go over to my Codex Agents. Um, I'm going to switch to let's do medium over here. Let's close this. Can I make this full screen? All right. Um, so let's give it give it a task. Um, spin up 20 sub agents to review all the sub agents. So, this is a very simple task.
All I'm asking Codex access to do the same task which I was showing before wherein I wanted to review all the different sub agent personas in this repo and you can see that you know, there's it it already figured out that there's like agents and skills and it's looking into it. There are 45 curated persona files and what what it's going to do is it's it's going to create 20 reviewers and it's going to give them all of those um Tamil files and then it's going to review those.
And you can see that there's two things which is quite interesting over here. Number one, Codex automatically decided that this is potentially uh a complex task. So, it automatically kick-started the plan mode which is what's active over here. So, you can see that it essentially came up with five tasks to solve this particular problem. You can explicitly invoke plan mode as well, but in this case it decided to do it on its own.
It's it's then partitioning all of these persona files and then it's going to spawn 20 sub agents very soon. Um I swear it's faster. But um So, now what it's doing is it's um Oh. Uh so, for some reason on my on my particular setup I have a cap on six num like six concurrent agent threads that can be run at the same time. We can fix that, but to go back up, what we can see over here is that it at least spin up six agents, which is my limit Uh for now and I can see all of those agents, you know, working over here.
I can quickly see like what Jason the agent over here is doing or Hume and so on and so forth and you can see that something to note here is that the the main Codex model over here the main Codex model over here essentially created a persona, right? And I'm not to start it double down and it it gave the exact files that this particular sub-agent should review, right? And and additionally, it also gave it some some insight on there's there's repo guidance in repo.md, in contributing.md, in skills and so on and so forth and it will sort of continue going down this this route for all the different sub-agents, right?
And what it does towards the end stage is that it will tear down all of these sub-agents when when they have gone through um when they have gone through their whole process of looking through all the tunnel files and so on and if I go back to my main thread um you can see that two of the agents are are still working um but eventually like it would collate all of this feedback that it that it has gotten from all of these individual sub-agents and you know, proceed.
Um now you can you can think of this this is like a very simple sort of explorer use case, right? But you can think of this from for example, a cybersecurity perspective wherein you have a get commit or you have a particular get repo and you want Codex to spin up and run multiple you know, vulnerability um Sorry, one sec. You want it to create multiple sort of you know, vulnerability analysis from different points of views or from different hypotheses and you wanted to sort of tackle the same diff or the same GitHub repo and try and come up with like a vulnerability map, right?
And this is something we actually use um quite a bit or I personally use quite a bit when I'm brainstorming a particular feature. I would just spin up multiple um Codex sub-agents to sort of look through how I would approach a problem, right? So, let's say I want to add a feature, I would ask Codex to create a plan for what are say five or six or 10 different ways that uh that a model um that a particular feature could be implemented and then I would quickly double down on like and ask Codex to um then create multiple sub-agents to get me some sort of understanding for um for these tasks.
Sorry, my watch was constantly vibrating. Um and and so that's like a um that's like a quick high-level overview of how sub-agents work. Uh by default, we ship three sub-agents, um three sub-agents personas. Um Let me quickly open. So, by default, we ship um three personas. One is like a default general-purpose fallback agent. Another is a worker, which is sort of execution-focused. So, this is something that you would use for um you know, when you want Codex to write a particular feature request uh or work on a particular feature.
Then there's explorer, which is the same one which we used uh before. And and then uh for for each of these, you can double down and create your own Codex um sub-agent personas like we saw before and we will create one right now. Um something to note is um, is that these particular sub agents, um, they like for each of these, you can define what model you want to use. You can define what reasoning effort do you want to use.
You can define what sandbox mode do you want to use and so on and so forth. Um, the reason why this is important is for a review agent, you would almost always 100% want to use the review agent in read-only mode. You would never want your review agent to execute anything, right? Um, for same reason for like a cybersecurity vulnerability uh, assignment, you would want your um, your sub agent to always be in read-only mode.
But for a for a um, for like a docs writer or for something which like, you know, creates um, docs for a particular feature that you've created or a bug report and so on, you do want to give it write access so that it can execute stuff and also create a um, create a bug report for it as well. Um, something to note is that you can also double down and give these um, sub agents, you know, more capabilities by giving them uh, MCP access.
So, you you can just give um, you know, let's say you can give a sub agent MCP access to Sentry so that it can look through all of your um, um, all of your reports over there or like one sub agent access to your linear um, you know, backlog so that it can um, it can interact with linear. It can uh, read through all the um, all the issues uh, added to you, triage them and so on and so forth. You can also give them skills.
Um, so really you can um, um, if you really want to, you can quite heavily customize this entire setup for your own um, for your own use case. So, let's open um, our Codex app again. You can see that it went through all of these sub agents. It created bunch of uh, other sub agents just to go through all of these and uh, it came up with these findings. Um, it's like based on read me, based on contributing uh, performance investigator um, is over privileged um, P1 has a sandbox mix uh, sorry, verifier has a sandbox mismatch.
Same for writer and so on and so forth. And so you can see that this is already quite useful um, and it saves you quite a bit bit of time to be able to go through all of these uh, individually or sequentially and so on and so forth. Um, now let's go back and see a bit more about custom sub agents. Um, so as I mentioned that we ship three um, sub agent personas, but at the same time you can create your own custom sub agents.
In fact, we do recommend creating your own sub agents or just ask your your Codex to look through your past sessions and create sub agents for you. Um, both of these scenarios work and um, work quite well. So, in in this particular case uh, you can see that we have a PR Explorer sub agent which um, reads your um, your code base, uses GPT 5.3 Codex Spark which is our um, research preview model text only um, deployed on Cerebras um, and is blazingly fast, is quite fit for this particular use case.
And we set sandbox to read only so we don't want the model to sort of execute and we give it certain uh, you know, ex- in instructions. So, in this case we say, "Hey, stay in the exploration mode, uh, trace the execution path, you know, um, don't propose any fixes and and and just like, you know, search through and and and figure out like what what what exactly do you want us to do." Um, now let's quickly try and try and um create a sub agent.
So let's say we want to do um docs researcher. In this case, what I what I typically do is to just go and ask um Hey Codex, can you create this sub agent and um for me? Uh here's here's its persona. Um and then let's see. And so what Codex is going to do because Codex is aware about um about how it works and you know, uh what it's supposed to um do and where it's supposed to place uh all of these things. Uh what it's going to do is it's going to create uh create a tunnel file for this docs reviewer.
And in this particular case, this is this uses the docs MCP server which we created um um from the DX team um which packages all the API references, all the docs, all the guides, all the you know, toolkits and so on and so forth. And uh it will add that as an MCP server so that every time we ask it um ask it a question about hey, like what's the best way to use GPT 5.4 with web sockets or what's the best way to use GPT real time with uh um with I don't know, pick your favorite way of using GPT real time.
And um and can you create a React plugin for this and so on and so forth. Um Uh it would be able to reference all of these things. So I'm going to let it do its thing and in the meantime um head back over to the slides. And so just to go back, sorry, one second. Um what you can do just to sort of invoke um you know, a particular sub agent is you can say um hey, can you reviewer sub agent and review each and every persona based on the developer's docs.
So, uh in this case, you can you can essentially like use the same particular um sub agent uh leverage it again and then ask it to do the particular task that you want to do. Now, what are some like interesting ways that you can use this is um imagine like you have like a long build process or you have a test process. You can have a sub agent which can run your test case locally. You can have a sub agent which can uh always make sure to um oh, I'm I'm being told that I don't have as much time.
Uh um you can have a sub agent which can uh pull the latest from uh from GitHub as soon as you do a pull. You can have a sub agent which can you know, quickly um pull all of the context from a linear issue and so on and so forth. So, really like you can you can you can do this for you can leverage this for a lot of um things. And the best thing that I like to do is to just ask Codex to look through my past sessions and recommend me certain automations, certain sub agents, and so on and so forth that I can use.
Cool. So, now we're at the at the bleeding edge. This is a bunch of stuff which we have shipped in the past and we haven't really made as much of a splash about. Um so um what we're going to do is we're just going to quickly go uh around and see like what each and every one of these um do and and how you can leverage them. So, first and foremost is guardian approvals. This is an experimental feature. You can activate it today by just going on slash slash experimental.
Um So, it would be something like um Codex. Hopefully, it works. And then you can look at um experimental and you can um in in my case, I already used guardian approvals and you can activate it this way. Um what guardian approval does is um all of us, including myself, at some point were um guilty of using YOLO mode all the time, which means that you by default give unfettered um access to your coding agent to do literally whatever the hell it wants, right?
And this by all means and measure is not safe. Um hence, we came up with something called guardian approval, which for each and every time Codex needs um a privilege needs to run a privileged task, let's say it is uh can I remove this particular directory? Can I run a uh server? Can I expose a particular file to um um to the internet? Whenever all of these things sort of pop up, what Codex will do is it will spin up a new sub agent, right?
Which will, based on a particular prompt, try and verify whether or not this is something which needs my human interruption or not. Um and in most cases, it doesn't need, you know, human interruption, so it will just say, "Hey, go on. Run this particular, you know, privileged tool or privileged task." and so on and so forth. And um this way, what we what we hope to do is we hope to reduce the human fatigue uh that comes by just, you know, always sort of having to approve, you know, do this task, do this, run this particular bash script, or run this, and and so on and so forth.
In um in principle, how would that look is um trying to see if there was um Okay, it doesn't show show it to me right now, but if I just in the interest of time, I'm going to ask uh hey, can you run the dev server? And I'm going to instead of full access mode, uh which for some reason again, I'm not able to uh click on. Let's Let's try and see um if it if it invokes um guardian approvals. Whilst this this works, I'm going to head over uh to the next step, which is hooks.
Hooks is also something which is experimental right now. We're we're trying 24/7 to try and make this uh a better experience. Uh currently, Codex supports three hooks. One is after each tool use, one is at the start of a session, and third is at the uh when you stop a session. What hooks allow you to do is it allows you to programmatically ask Codex to do a thing X uh based on a particular event. So, let's say that when you start your um your Codex session, you want Codex to pull the latest from your GitHub repo.
So, in that in in that particular case, you would want to set up a start hook. Um if you want Codex to do something after each tool use, let's say um for a lot of researchers who want to document each and every tool use, they might have like a per tool use hook where in they document what Codex has done Uh per session and so on and so forth. So you can do that with that. Um And last but not the least, something which I personally use is the stop hook, which is when I'm running long-running tasks, I would at the end of each turn of Codex, I would ask it to keep going.
So that like it just continuously you know, continuously keeps running a particular task. And um In in theory, how this would look like is um is Where is it? is Sorry, one second. Wow, I was really prepared for having more time. Um I have to say. But um in theory, how this would look like is um is that you have some sort of a Python script um and you have you define like a hooks.json. So in this particular case, you can see over here that you have a pre-tool use um you have some sort of a you know, matcher.
You say like on startup or resume, run this particular session.sessionstart.py and so on. Uh and you can define how you want to uh in this particular case. Um So what I did for for example uh for the sales dashboard example that I've been showing you so far is I created a hook for stop, which runs um this Python script which is keep going.py, um which is every time it encounters the stop um um hook, it would just ask Codex um to keep going, do one more pass, run one one solid validating command, type in one more thing, and then stop and give the result.
And so for really long-running tasks, you can just set it up and like ask it to continue doing its own thing. Um Last but not the least, um we have personality changes, which means that you can go on Codex and you can ask it to um quickly look at personalization. You can set up different personalities. You can set up a more friendly personality or a pragmatic personality based on whatever you want to do. You can also add custom instructions, so you can ask it to always cite whatever it is it is doing and so on.
Um Right. And then last two things um is we released something called Codex Security. This is our state-of-the-art model, which allows you to find and fix vulnerabilities in in your GitHub projects. And you know, essentially what it does is it would go through commit by commit and it would create a vulnerability patch and then and it would use Codex to then sort of patch the same changes as well. Um lastly, as I mentioned before, we released a cloud code plugin, which allows you to use Codex in in cloud code.
This is something which was surprisingly used quite a bit by the community. And this is something which allows you to sort of ask Codex to review whatever it is that you've done so far, run an adversary adversarial review, or just like ask Codex to rescue whatever changes you've done so far as well. Um that's it. Thank you so much for for joining us and feel free to ask any questions that you might have. Hi. So, we don't have a lot of time for Q&A.
Unfortunately, we should have started maybe a little bit earlier, but happy to take maybe a couple questions in the room and then we'll stay here anyway, so if you have questions and you don't have anywhere to be, you can come to us. Yeah. >> Thank you so much. And I have a question you said a couple of times that there's like a way to scan that Codex scan past sessions and basically give >> Yeah, so what you typically do is like all of the sessions within codex are put in dot sessions within a particular within the same dot codex folder and codex has the ability to just like scan through all your sessions and then you know do things.
Yeah. You can use it you can use codex app you can use codex CLI anything you just have to ask it to look through the sessions and do whatever you want to do. Nice. There's another oh okay maybe a couple more. Yeah. In the back here. Hi. Hi. Is there a way to hand off a task to a cloud agent? So let's say I'm here working on a task and I'm have to close my laptop so I have to a cloud agent. >> Yes definitely we didn't touch on that but actually you can do that from the codex app directly like maybe you can you can show your screen but you can either work locally and as you mentioned you can do it like we support get work trees as well but you can also just select cloud here and you can select the number of times this task should run like you can parallelize we call that like best of n so you can like run it four times in the cloud and then just pick the best output so that's something that is like built in in the the codex app in the ID extension and you can also like access it directly from the the web interface.
And there's more cool stuff coming on that very soon. More what? There's more cool stuff coming on that very very soon. I think there was one right here. Yeah. Thank you so much. My question was actually about the cloud UI as well cuz today Sabatians aren't supported if I'm not wrong. And especially the thing that bothers me is it doesn't use the the skills that are in the repo. Is that coming soon or So there's like a at the risk of you know, talking about the whole road map.
We we we definitely have a lot more changes coming up on that particular front. I'm not sure if skills within cloud is going to be as soon as I say that it's going to be but it's definitely at the top of the mind and we do want to sort of add you know, give you the ability to sort of like have your own trusted MCP servers to be able to run there or CLI's and so on. And also the ability to just like have SSH agents that you can just spawn off a particular task to on a VM and so on.
So lots of work on that like It can use skills in the repo, right? That that's checked in. It's not on cloud tasks. Yeah. But like if you like it it reads instructions and stuff and they can like find it and like still see it since it's in the code base. It's more like the the skills that you have locally that would be the same. The reason why we don't allow it on on cloud is because there's no way for the sandbox to know whether or not a skill is trusted or not, right?
And so that's why we we we don't allow it. A skill can package like a Python script or an or an executable. It won't execute things but like if you have, you know, like things like resources, it can access it technically because it is like in the repo. It's just yeah, it's not as good So I have to request it. Yeah. Yeah. Thank you. Thank you. Were there other questions? Cool. Have a great day. Enjoy the day. And if you have any other questions, we're going to be around today, tomorrow, and also maybe on Friday.
Feel free to reach out or just like drop a DM and enjoy. Thank you.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.