Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Cole Medin · @ColeMedin
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in Cole Medin's most watched videos.
Most replayed moment at 10:58
3.0x that video's typical replay level
cache is performing in production with real user data. And the best part is better DB is open source and free to get started. So I'll have a link in the description. I'd highly recommend them as a tool to help you scale manage your costs for agents you're deploying to production. And so now Google is saying with
Said at 10:50
Most replayed moment at 3:20
3.2x that video's typical replay level
doesn't end up becoming the standard down the line for personal agents. There's going to be something like this. And so it's good to understand this now. Okay. Now, let's really get into OKF. So there are two things that they're standardizing here. The first is how we are organizing information like our
Said at 3:13
Most replayed moment at 9:16
3.5x that video's typical replay level
not extremely difficult to get all this set up like it used to be. And the best part is the agency CLI is free and open source. You can take these skills, bring it into any coding agent, and see how easy it is right now to build any AI agent. I'll have a link in the description. I'd highly recommend
Said at 9:09
The graph counts replays. It does not show where viewers stopped watching.
Words
3,639
Runtime
17:13
Speaking pace
211wpm
Reading time
15min
211 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
So, we have a new AI model that was just released within the last week called Jev. But, Jev is special. Jev is not just another large language model. It introduces an entirely new class of AI models called system one models, which are master decision-makers. Now, that might not sound super exciting at face value, but having a master decision-maker is actually incredibly useful, especially because of how fast and cheap Jev is. We'll of course talk about that as well. And I'm genuinely excited for this. Like, I haven't been excited for an AI model for a while now. I got to be honest,
106 words, the words spoken in the first 30 seconds at 211 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 212 |
| Average words per sentence | 17.2 |
| Longest sentence | 53 words |
| Questions asked | 10 |
| Sentences containing a number | 16 |
Most used terms
Filler phrases
74 in total: like 39 · kind of 14 · actually 11 · you know 3 · I mean 2 · right? 2 · basically 1 · literally 1 · uh 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
So, we have a new AI model that was just released within the last week called Jev. But, Jev is special. Jev is not just another large language model. It introduces an entirely new class of AI models called system one models, which are master decision-makers. Now, that might not sound super exciting at face value, but having a master decision-maker is actually incredibly useful, especially because of how fast and cheap Jev is.
We'll of course talk about that as well. And I'm genuinely excited for this. Like, I haven't been excited for an AI model for a while now. I got to be honest, I've been kind of burnt out with all the new LLMs coming out. Cuz every single time we have a new model, there's always the rush to incorporate it in our workflows. There's all the hype on the internet for "Look what this does for your second brain. Look what kind of beautiful websites you can make.
Look at these beautiful scenes that I'm generating in Blender." And you see that time and time again. It's just like, man, I want something new. And we have something new and genuinely innovative here with Jev. Now, Jev has been out for a little bit, so you might have already seen a video or two on it. But, I've specifically waited to make content on Jev cuz I wanted to actually build with it as I tried out initially before I go and just talk about it.
And so, of course, we'll start with a high-level overview of how Jev works, its limitations, and the really incredible benefits. But, then I want to get into how I've been building with Jev and show you some really cool things I've been using with it already. And then, there's one other thing I want to hit on, which is there's actually a good amount of criticism for Jev as well. The idea that it's not really truly innovative and it's more just taking ideas that already exist in the industry for classification models.
It's kind of true. It's something interesting to talk about. And so, we'll hit on that, as well. So, the most important thing to understand with Jev is it is not a large language model. You can't have a conversation with it like you can with ChatGPT, Claude, Gemini, because it doesn't generate text. It's sole responsibility is to take a bunch of input around a situation and generate a decision. And so with this, there's a brand new training algorithm that they called RLCD, reinforcement learning for calibrated decisions.
This is different from RLHF or reinforcement learning with human feedback, which is the algorithm used to train generative AI models. And so the promise with Jev and system one models is that we are solving a lot of reliability and hallucination issues we see with generative AI specifically for making decisions. And I love what they say here, "Our claim may sound too good to be true, but the bitterest lesson in AI is that optimizing for the right task gets you an unfair advantage." And I agree with this 100%.
I love how they sign off here, "May your intelligence be ever reliable." And if you follow my channel, you know that that's a really important thing for me that I'm always chasing. Like when I'm building Arkon and my AI coding workflows, I'm all about adding as much determinism in these workflows as we possibly can. And Jev is another tool in our tool belt to do that. Now, before I go and explain more with Jev, I want to show you what it looks like in practice so you can see what the inputs and outputs look like.
It's quite different from a large language model because it's never going to be free-form text. It's always going to be a decision. In fact, it looks a lot like structured output that we have with LLMs. Though there are a lot of differences and we'll talk about that as well. So the best way to describe the input to Jev is it's really a two-part input. You have the situation that it is analyzing and then you have a set of questions have a multiple choice answer.
It is never generating free-form text. Remember, it is always making a decision based on the situation. And so for this basic example right here, you can imagine Jev is integrated into some kind of customer support agent. And so this person, they're having trouble connecting their Stripe account. They're losing sales and they need help ASAP. And so we're asking Jev a series of questions that it's going to answer in parallel and it's always going to pick from the criteria that we give it, multiple choice.
And so, the output quite simply here is going to be its answer, its choice for each one of the questions, as well as its confidence score. And so, here it made a routing decision. It's 64% confident that it should go to the billing department, this complaint. And for the frustration, a score of one, that means that the customer is deemed frustrated but civil. So, doing sentiment analysis as well. Okay, so that's cool, Cole, but large language models can already do this.
They can already output structured JSON to make decisions in And yes, that is true, but there are three massive benefits to Jev. It's the perfect trifecta. Jev is faster, more cost-effective, and more reliable when it comes to making decisions compared to large language models. At least what they claim, and what I've seen, it has a 0% failure rate for structured output. So, it never has any kind of malformed JSON that would make a future step in your workflow fail.
And Jev is incredibly cost-effective. This graph right here is logarithmic. And so, Jev is dozens, even hundreds of times more cost-effective than all of the best large language models. There are a couple of models that do actually make better decisions if you give them the time than Jev, but it is incredibly cheap. Like, look at the cost per million tokens of Jev compared to something like GPT-6 Astra or Claude Fable 5.1.
And so, if you combine the incredible cost-effectiveness with the speed that we have with Jev, it's able to make decisions 20 to 200 times faster. It's 40 to 1,000 times cheaper. That together means that you can build these systems having a lot of intelligence, making snap decisions extremely fast, and even making hundreds or thousands of decisions in parallel. So, there are a million different use cases for Jev that are super powerful, very practical.
But the coolest one that I've been working with right now, I want to show you just really quickly, is using Jev to play video games exactly as a user would, making decisions as quickly as us or even faster. And so, I've been experimenting a lot with using Jev within my AI software factory to test out games as I'm building them as the large language model is actually writing the code. And so, you can see in real time it's making all these decisions with confidence scores for all the different actions that I'm proposing for it, and it seems like I'm playing.
It's attacking, moving to enemies, dodging, but I have hands-off the keyboard. This is so cool to watch. And I've tried to get large language models to do this kind of thing, but it just doesn't work because they can't process things fast enough, and it would be way too expensive. So, this by the way is actually my game. This is running on localhost right now. This isn't just some like Twitter demo that I have up, though I'm going to show some of those as well.
The important thing here is that Jev can't create this game, but it can definitely interact with it in a way that a large language model never could. So, it's not like you're going to totally swap all your large language models for Jev right now. It doesn't work that way. You're just going to put Jev in your automations where you have those decision points or any kind of classification step. And so, the best workflows for AI coding or any kind of business use case going forward is going to be a combination of Jev for the decision-making and LLMs for the other reasoning.
The sponsor of today's video is Firecrawl. Every AI agent that I build eventually needs access to the web, and there are a lot of agents out there that have these capabilities out of the box like Claude Code, but if I'm building my own AI agent with Pydantic AI or LangGraph or Pie, I have none of that. And even if you are using Claude Code, the search capabilities built right in are very inefficient and token-heavy if you haven't noticed before.
And Firecrawl has the solution for this. They call it the context API for AI agents. That's exactly how I use it. And they have an MCP server that makes it extremely easy to bring their context API into any AI coding assistant or other AI agent. So, for example, with Claude Code here, I just copy this command, go into a new terminal, paste it in, just a single line to get the MCP server added. So, now when I go into Claude for the first time, I simply have to do {slash} MCP to then set up the authentication, and then I'm good to go.
So, I'm showing you the full flow here in just like 20 seconds. Authorize, and then back over to the terminal, authentication is successful, and I can now start sending in my requests. And the search capabilities of Firecrawl are powerful. It's not just a Google search. It's able to directly generate queries that answer my question instead of just performing a really broad web search like you'd usually see in something like Claude code.
So, right here I asked, "What are people running into when upgrading to Pydantic AI version two?" A specific but powerful example cuz it has to look through a lot of context to answer this, but it's able to do so in only three calls to the Firecrawl MCP server. So, Firecrawl gives me exactly the context I need. It can also give me the full page as clean markdown. The MCP server is easy to use anywhere, and they also have an SDK if we want to build Firecrawl directly into our custom agent tools.
So, it's super easy to use whether you're building your own agent or using something out of the box like a coding agent. And Firecrawl is free to get started with a thousand credits a month, and you don't even need an API key to use their MCP server. I'll have a link to them in the description. Now, of course, I've been doing a lot of testing with this myself, building larger AI coding workflows with my open-source tool Arkon, combining Jev with LLMs, using the right model for the right step.
And so, with this Jev PR triage workflow, essentially what it does is we have classification and routing at the start that figures out what kind of review we need to perform on a pull request, and then go and do that review. And of course, for the first two steps here, classification and routing, I'm going to be using Jev because we're just making decisions here. And so, it's very cost-effective, this workflow, because of course, Jev itself is cost-effective, but then also we get to decide what kind of review we're doing cuz we don't always need a super deep AI review on every single pull request.
And so, this is just one simple example of the lot of testing that I've been doing with Arkon. So, also, let me know in the comments if you want me to make more content on this. I'm definitely going to continue to explore using Jev within AI coding workflows for testing things like I showed you with my game, for classification steps with things like issue triaging and pull request review. The possibilities are endless.
And if you've been following my channel in Arkon and you're curious, this is the exact workflow that you just saw in the Arkon UI. So, I just simply call this Python script that makes the classification with Jev. And so, for my Jev usage, I'm going directly through Open Router. So, Open Router was super fast to make the Jev model available along with all the other LLMs that you can use there. And then also, if you want, you can go directly through typesafe.ai.
Typesafe is the company that created Jev. And so, of course, I'll link to this in the description. By the way, they're not sponsoring this video at all. I am genuinely excited for this model. I hope that you are too, just going through some of the use cases with me here. And if you're not sold yet, let me show you some more use cases. So, this is some experimentation I've been doing myself. I've seen a lot of other people on the internet do this as well, using Jev as an LLM router.
It's a really common use case where you have a bunch of different LLMs that you want to pass the right requests to, right? Like sometimes, for the sake of cost, the simpler requests go to the faster model. Ones that require deeper reasoning you want to send to the strong model. Traditionally, you've used yet another LLM to make the routing decision. But again, with Jev, even with tiny LLMs, it is going to be faster and cheaper, and of course, more reliable.
So, the situation we give as input to Jev is the query that we want to route. And then, the multiple choice that it has to answer is which one of these models should we route it to? The strong, the coding, the open, or the fast? And this is just a quick visualization I put together to show you all the testing that I've been doing. But for a deeper question, it routes to the strong model with a confidence of 100%. Uh this, you know, convert this bash to PowerShell, a little bit of a coding task.
It has a 98% confidence going to the coding model. And some faster ones here, like convert 72° F to °C. Yeah, this can definitely be handled by a cheaper model, like GPT-5.6 Luna is the faster one here in Open Router compared to, I don't know, like what's the strong one here? Yeah, Claude Sonnet 5, for example. Now, these numbers aren't the best. This is just a really small subset of all the testing that I've been doing, but the really cool thing to show you here is that out of the dozens of the tests that I have visualized here, it costed me 4/10 of a penny to do all of this routing with Jev.
And the average time it took was 2/10 of a second to make each one of these routing decisions. Super cool. Okay, so that's enough of my testing. I hope you liked it, but let me show you really quickly what other people have been sharing on the internet as well. So, this person on X posted using Jev to play the classic game Doom. And it actually looks a lot like my own testing with my own game, where we have the decisions that are being displayed on one side right here in real time as Jev is playing the game.
And then, of course, the game itself. It's so cool how it's able to play something like this. And yeah, it's a basic game because you can't just give it like millions of decisions, but this is still incredibly impressive. And of course, I'll link to all of these resources in the description. Another really powerful use case for Jev that you probably thought of at this point is using Jev for browser automation. So, more traditional tools like Playwrights and Vercel's AI agent browser CLI, it's always driven by an LLM.
And I use these every single day as I'm building web apps, full stack apps, but it's always slow, right? Like the slowest part of my AI coding workflow for any kind of full stack app is always when it has to validate things visually and navigate the browser. But now we don't need an LLM to do it. We can use Jev because every single situation is the current layout of the site, and it just has to decide with multiple choice the next action to take like click this button or type in this input.
And so, this is just one really cool open source repo that I've seen. There's probably going to be a lot of Jev browser use tools released in the next couple of weeks, but this is one of them. It works incredibly well. And then I also found this really cool visualization of Jev playing Pong. And so that's the top row right here. The game is slowed down basically to the rate that the model can handle. And Jev can pretty much handle it at human rate.
And then other LLMs down here like 3.8 flash, Claude Haiku 4.5, you can see the game has to be slowed down a lot for it to actually process where the paddle needs to be as the ball is coming. And then one last resource I want to show really quick is this open source repo that curates a list of projects and just general use cases for Jev. So, I'll scroll down in the read me to current coverage. We got classification and routing.
Of course, that's going to be the most common one. I mean that's literally what Jev is made for. But then using that for agent decisions, verifications and guardrails, calibration and research, games and simulation, finance and trading. There are so many cool use cases to poke around here. So yeah, I'll link to it in the description. Just check this out. You just get your imagination going here as you go through these different use cases like I'm trying to do for you in this video.
Cuz when you you think of Jev as this glorified classification model, it's not very exciting at first. But once you realize what you can do with it, man, it's the world becomes your oyster here. And speaking of Jev being a glorified classification model, that's the last thing I want to hit on really quickly cuz it's the biggest criticism that I've seen for Jev and I've actually seen it quite a bit. A lot of people say that we've had the idea of a classification model in the AI industry for decades.
And so we're just reinventing the wheel here. And that's it's true to an extent, especially when we use it for very basic things like this example I have in the type safe playground. But what really makes Jev powerful is how general it is. Like I think the best way to describe it is it feels like there's still the intelligence of a large language model operating behind the scenes producing the structured output answering our different questions.
Like I just had this really silly one right here. Which country has the coolest buildings? I gave it a few options and it's actually says Germany with 63% confidence. A very opinionated thing. I mean, who knows what it's going off of. But if I run this over and over and over again, the numbers change a little bit, but it actually always says Germany. Which is interesting. Very cool. So anyway, anyway, I built a lot of classification models in the past, but it's always for a very specific task and you have a very specific data set.
So I've used, you know, like TensorFlow and PyTorch to build classification where you give it a chess position and it says who is winning or it looks at an animal and identifies which animal it is. But then if you have it do a different task, it can't do it at all because it's so specific. But with Jev, it's classification in the general sense. You can give it any kind of situation for customer support or opinions on countries if you really want and it's able to give you a response here.
Now, this is kind of a silly example, but for anything more objective, like what kind of pull request review level does this need? Or what model should we route to here? Or what's the next best action in this video game? Like Jev can just handle any of that and it's so accurate. So I hope that you found this interesting. All the super cool use cases for Jev. I would encourage you to try this right now either through OpenRouter.
They also have a waitlist that I was able to get into within a day. I'll link to that in the description as well. And I will certainly be doing a lot more content, especially with Arkon and how I'm using it in my AI coding workflows. So stay tuned for that. So with that, if you appreciated this video and you're looking forward to more things on AI coding and Jev and system one models, I would really appreciate a like and a subscribe.
And with that, I will see you in the next video.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.