Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Simon Scrapes · @simonscrapes
Words
8,870
Runtime
37:22
Speaking pace
237wpm
Reading time
37min
237 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
The most talked about AI model this month can't even write a sentence. You can't chat with it, and it can't explain itself. And this is what people built with it in the first week it was released. We've got a browser that fills in Google Flights in about 7 seconds. And if you've ever tried to use Claude with a browser, by the way, you'll know that it's incredibly slow. We've got 3 million website sessions capturing lost customers, abandoned checkouts, and more sorted in 40 seconds for about $2. And then we've got someone else driving Chrome with their voice, where it even takes actions before they finished their full sentence. So, it's called Jev, and I've
119 words, the words spoken in the first 30 seconds at 237 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 525 |
| Average words per sentence | 16.9 |
| Longest sentence | 61 words |
| Questions asked | 48 |
| Sentences containing a number | 59 |
Most used terms
Filler phrases
130 in total: actually 56 · like 44 · basically 14 · right? 5 · you know 5 · kind of 4 · literally 1 · uh 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
The most talked about AI model this month can't even write a sentence. You can't chat with it, and it can't explain itself. And this is what people built with it in the first week it was released. We've got a browser that fills in Google Flights in about 7 seconds. And if you've ever tried to use Claude with a browser, by the way, you'll know that it's incredibly slow. We've got 3 million website sessions capturing lost customers, abandoned checkouts, and more sorted in 40 seconds for about $2.
And then we've got someone else driving Chrome with their voice, where it even takes actions before they finished their full sentence. So, it's called Jev, and I've been through all the official docs so that you don't have to. And by the end of this, you'll know what it is, why it matters if you use Claude already, how to run it from Claude code and ask it the right questions, as well as how to get the most out of it.
So, let's start by talking about what's different about Jev, because some of this isn't new. So, every piece of software you already use is already full of if statements. For example, if an order is over $10,000, we're going to flag it for review. The catch, though, is that an if statement can only check things a computer can measure. So, we're comparing it to a number, a date, a tick box, a criteria. But, it can't check whether an email sounds angry or whether an order looks dodgy, because those aren't numbers.
They aren't checkable things. They're not deterministic. So, Jev is doing exactly that. It's turning those messy inputs from customers into numbers. So, the if statement then becomes, "If this order looks suspicious, okay, we're 95% sure, flag it for review." So, now it can be read, and a decision can be made by a computer. And we'll talk about how you determine what counts as suspicious, because we're going to use Claude and you to write the questions, and you are going to set the criteria for yes, the 95% confidence level, for example.
So, we're going to send a bunch of questions to Jev. Jev is going to answer each one with how sure it is, and your code only flags the order when all of them come back above that threshold that you've chosen, so that 95%. And this might feel a little bit different to your traditional LLMs, like your Anthropic chat models. And that's why Type-Safe, the company behind Jev, call it a judgment model. So, it was built by a founder, Diogo, who helped build the methods behind ChatGPT at OpenAI.
So, this is a massive announcement. And if you compare it to something like Claude, on the surface, you can absolutely ask Claude, "Does this order email look suspicious?" And you're going to get a text response from Claude like, "You're absolutely right. This does look suspicious. Would you like me to draft an email to the customer?" It's going to take about 10 to 15 seconds to get to that point. Versus Jev, on the other hand, who's given a bunch of questions to look for up from.
The answers to those questions are going to determine how suspicious it actually is. Like, is there an address mismatch between the billing address and the delivery address? And then it's given the information it's actually processing. This is called the state, and in this case it would just be the individual order email that we're assessing. And then the response from Jev is not going to be a text response. It's only ever going to give you a number from the options that you've defined.
And the number is going to tell you the likelihood of it being suspicious against the criteria that you defined. So, it's never going to write an output like Claude. And there are certain advantages to this that we'll cover in a moment. And this is totally different from the chat interfaces that we're used to when we see things like LLMs, large language models. This is closer to these deterministic machine learning models that is another style or type of AI.
And Type Safe specifically call this what they call a system one model, after the book Thinking Fast and Slow. And system one is basically a quick name for a fast gut call. Where system two is sitting down and working out the reasoning, right? So, Claude, ChatGPT, and every other reasoning model are system two, but Jev sits in system one. So, why is it so special then compared to everything else that we've seen today with our traditional LLM chat models?
So, there's a few reasons why not writing as the output is advantageous in some use cases. Firstly, it's cheap and extremely fast. So, it's suggested that it's 42,000 times cheaper than traditional LLMs and 20 to 400 times faster. It's about 4 cents per million tokens in, and the output is practically free. The second reason it's advantageous is it can't make up an answer. So, if you give it three options, it comes back with one of those three options every time.
So, this is actually a pro and con, which we'll come to later in the specific rules that you need to give Jev when you're operating with it. The third advantage then is it that gives the same answer or almost the same answer twice in a row. So, it's much more consistent than LLMs in answering the same question over the same data set. So, the same test, they ask the same questions over and over, and the answers or the output from Jev moved less between runs than any of the chat models they tried.
Even with the prompt set to be as repeatable as possible with the other models. And then the entire point of this is the number that Jev gives you as the output makes something that's usually uncertain, like a fuzzy customer email that's messy, it's got lots of text, and it makes it actually usable by software. So, Type Safe have trained this model so that the probabilities are calibrated. So, in plain English, it basically means it works like a weather forecast.
So, when Jev says 0.8 is the output, it should be right about eight times out of 10. And based on that eight times out of 10 or the confidence in the output, that's what lets the software decide whether to act on it or not, whether to route that query somewhere else. The next actions are decided based on that number, and Jev gives you the probability of that specific question being true or false. And we'll see the different types of options we're able to give to Jev in a little while.
Compare that to a chat LLM model which are trained to sound right to a specific person, and that's why we have that bias where it agrees with everything we said. Jev is not built in the same way. Now, super important, if you already use Claude or something like GPT, it does not replace it. It's going to work alongside it, and it's going to add advantages to certain workflows. So, Jev replaces one entire category of work that Claude does today.
And whenever I say Claude, you can, you know, replace that with any of the LLM providers that you want here. And that category of work is going to be the fast decisions that we need to make in bulk inside software. So, the code that you're going to pass to Jev hands Jev that information, like our customer email, and a set of tightly defined questions which we want to know the answer to, and Jev is going to send back a probability, a category, or a score.
And then your code, off the back of that, is going to use that to route it, to rank it in order, to flag it to a human, or to carry on. So, it's going to take actions off the back of that potentially. But, everything that needs a conversation, a chat back and forth, or reasoning, or writing is still going to stay within our Claude interface. And Diogo, the founder, specifically said on the Latent Space podcast that Claude is the conversation, and Jev is the plumbing underneath that.
And it's important to actually reflect on the fact it cannot do everything, and in fact, it can do very few things. But, the things that it can do, it can do them really, really well. So, we mentioned earlier that each question it answers can be split into one of three types. Now, let's run through each of those types using that angry customer example again for ease. So, first, we have yes-no questions. These are called a null, and this might be, "Does this customer sound angry?" Let's say Jev gives us a 0.9 back.
So, it reckons it's a 90% chance of yes, this customer is angry. And then your code, your script, important to note this isn't Jev, this isn't your LLM, it's a script that we can get Claude to write, then does the rest with a normal if statement. So, if the result is above, let's say 0.8, we're going to escalate that to a manager. So, the manager can then directly respond to that customer. Or if it's below 0.2, i.e., they're definitely not angry, or Jev has deemed them definitely not angry, then we're not going to escalate it, we're going to deprioritize it.
But, it's important to understand that Jev never actually escalates anything, it just gives you the number, which represents the probability of it being a yes in this case. Where 0.9 is a confident yes, the customer is angry, and 0.5 means it has no idea, it's like a coin flip. And then the code or the script is going to take the decision based on that number it gets back. The second type we can give it is called a choice.
It's exactly as it sounds. You give it categories that you predefine, and in the angry customer case, it's like, "Which team should handle this query? Should it be billing, technical, sales?" And it's going to come back with a probability for every single one of those options. It's not going to say yes, it's billing. Yes, it's sales. It's actually going to say billing was an 85% chance of being correct. Technical was a 10% chance.
Sales with a 5% chance. And the code off the back of those results is then going to route the email to the winner. So in this case it's going to be billing. So you're going to use choice when you're routing the output somewhere. And when the options do have an order, we're going to use the third type, which we'll talk about in a second. Now because we've got more than one option there, we've got billing, technical, and sales, you actually get a second number returned by Jev called confidence.
And that's worked out from how spread out those probabilities are. So the 85% for billing is confident. But let's say billing was at 40, technical at 35, sales at 25, then we're less confident about that or Jev is less confident. Billing is still going to be the chosen one, but the confidence is going to come back low. And that's the number your code is going to check before it trusts that pick in the first place. So it's going to check that confidence number too.
And that's exactly what the browser demo that we showed at the start is doing. Every button and box on that page for the flight booking is going to get a number. And Jev answers two choices at once. What to do with it? So do we click, type, or scroll based on the action or the input from the voice? And which number to do it to? And we've basically ordered everything on the page to get a number. And we're going to pick element three or four or seven.
And there's no order to either of those. And so we know that this type is going to be a choice. And the confidence then is going to be that safety catch. So if Jev is confident, i.e. puts it above that confidence level, then the click is going to happen. With below it, the agent is going to stop and ask the user. And you can raise or lower the threshold for confidence and what is acceptable based on a specific action.
If it's something that's going to be detrimental or highly costly, then we need a higher confidence level to permit an action or a higher safety catch effectively. And the third one then is a score. So it's the same idea as choice, but the options have an order to them. And you describe each one in words. So the question might be how angry is this customer? And the results are going to be a score. And the score is going to be between calm, annoyed, and furious.
Those are the three options. So Jev is going to come back with a position on that scale between calm, annoyed, and furious. And it can actually land between the two levels. So, 2.4 out of 3 means they're mostly annoyed because that was the second option, but they're edging towards furious. So, when we're sorting an entire inbox of customer service emails, we can actually rank them. And this is what our code does or our script.
The inbox by the worst first. So, the three out of threes on top, the null out of threes or one out of threes right at the bottom. So, you'd use the score then on anything you'd put on a scale. This is why it's important between choice and score that there is an order to things here. So, it's like how urgent is this ticket? How good a fit is this lead for what we actually sell? How senior is this candidate on this scale?
How bad is this bug? Is it cosmetic bug? Is it an annoying bug? Or is it a blocking bug for somebody taking an action on our page? So, use it when you want to rank or sort and not root a response necessarily like choice. And because it's got several different levels, it also comes with a confidence the same as choice does, but null does not come with a confidence. It's either yes or no. So, the website session demo I showed at the start of the video is mostly that first type, a yes or no per event.
Was this person rage clicking? Was this person dead clicking? Did something error on the page? Then we actually add this score as a layer on top of that. So, how bad was this whole session for this user? So, that when we're sorting those sessions, we can come back and actually sort them from worst first, and a person then can only open up the top 10. Now, you can of course do this whole process in a large language model, but it's not optimized for making these probabilistic decisions, and therefore we can do an astounding amount of them in a very short space of time for very, very cheap.
And the way to choose between these three different methods, the null, the choice, or the score, is working out what your code does with the answer afterwards. So, if it needs to act or not act based on the answer, that's going to be a yes or no or a null. If it needs to root the response somewhere, that's going to be a choice. Or if we need to sort or rank by it, then that's going to be a score. And most importantly, all three of these types can actually go into one request to Jev.
So, here's an example of one email going in with all three questions attached, and this is what comes back. So, we've got a null first, a question that says, "Does this customer sound angry?" We've got the null 0.9. So, basically it's saying for the code to escalate this to a member of staff. Then we've got the choice, which team is this? And it says, "Billing is 85 likely, technical 10, sales five." And we've got a high confidence level.
So, it's going to route to the billing team. And then the score, how calm, angry, annoyed, or furious on a scale of four here. And we've got a 2.4 with confidence of 80% or 0.8. So, it's going to go near the top of the pile for the emails that the billing team are going to answer. And they're only going to see the ones where the customer actually sounds angry. So, they've got a short list of answers because Jev has actually processed a huge amount that don't need to go to a customer service rep.
So, we've got one email in, but three answers out in a very, very short space of time. So, now let's move on to how to actually pair this with Claude code so that you can use both tools to their full advantages. But first, I'm about 1,000 subs from that magic 100k and that plaque on my wall. So, if you're enjoying this so far, make sure to sub down below. So, anyway, there's two ways to pair it with Claude to get the best out of both tools.
And Type Safe are really clear in their docs about what it's bad at and therefore where you should use Claude. And that's reasoning, specialist knowledge, and anything that effectively needs writing. So, the first way in which you can pair it with Claude is actually using it around Claude's features for checking things. So, before Claude reads something, you're going to get Jev to decide whether Claude should actually read it, whether you should spend tokens on that thing.
So, Type Safe use a guardrails example, which screens every single message that comes into Claude for jailbreaks and harmful requests on the way into Claude. It's then going to screen Claude's reply on the way out before the user sees it. So, that's a good example of Claude in the middle, but using Jev as guardrails on the before and after. And you can also use the same idea after the LLM processing. So, in their citation example, an LLM wrote a document with eight citations in it, and Jev checked each quote against the source.
Four were found to be fine, one quote was a non-existent citation, and one said the actual opposite, and two were actually rooted to a person because Jev wasn't sure. So, that is like post-processing of Claude and LLM content gone through Jev to make sure that everything they were putting out was factually accurate. And the second way you can use it with Claude is actually using it instead of Claude for the bulk work, for the tasks that are larger in scale, but also very deterministic.
So, don't use Claude to sort 5,000 rows, tagging every ticket in your inbox, scoring every single lead. Claude should only get the rows that Jev wasn't actually sure about or the ones that need something written. If it's that deterministic scoring, then we should use Jev instead of Claude for that stage in our workflow. And this is the one that you've probably seen all over X or Twitter. Here's someone scoring 700 sales leads in 40 seconds for around 9 cents.
So, something that is bulk work would take significantly longer using Claude and also cost a lot more, too. And actually both of those examples we're going to build together in the next 5 or 10 minutes using a customer service example. So, let's get that set up inside Claude Code now. So, Type Safe actually shipped a skill for use in LLMs like Claude Code or harnesses like Claude Code, and it's two lines to install. So, you can either do this through the console once you've signed up on typesafe.ai, you can copy this agent prompt, or you can actually head to the docs.typesafe.ai/agent-skill, and you're going to run these two commands in your terminal to get it inside to Claude Code.
And you can also, you know, take these commands directly from here or use it with other agents there, too. I recommend to actually use this that you need to make sure that you name the skill in your prompt. So, when you're sending a command to Jev, for example, we're going to say use the Type Safe skill, and you'll see that in all of our prompts. So, we're going to run those two commands. It's going to add the marketplace.
We're then going to reload to make sure that the plugin's installed. And you can see that that's now a success. So, once we reload or load up a new session with Claude Code, we can run /plugins, and we can see, if we go to installed, that we've got the Type Safe plugin now enabled. You're going to head into the API key section, create a new key, and you're going to put that in a .env file under the name Type Safe API key.
Then to check it's working, we're going to say use the Type Safe skill. That's really important. Check my key works. Ask Jev one yes or no question about and this is our input as if it's a customer email, right? We've been charged twice. Please fix this ASAP. And we're going to ask Jev to assess it on the question, the customer is reporting a billing problem, and to tell us the raw answer from Jev that we get back. So, with this, we should just get a yes or no in the form of a number that that deems is this a billing problem or is it likely to be a billing problem?
Now, before we run the output, if you're not working in Claude code directly, and you actually want to work in the desktop app, I haven't tested this yet, but it should work exactly the same. All you need to do is go to find the skill.md on GitHub. You're going to open this page and then effectively download and upload the skills in the customize section inside your skills. So, you're going to add upload the file and then you should also have access inside the Claude desktop app to using Jev as well.
So, let's run this prompt. We're going to use the type-safe skill and let's see what comes back. So, you can see here, this is the request based on the type-safe skill that it's actually determined. So, the state is the information we're sending, charged twice, please fix this ASAP. That's the email we're sending it to Jev latest. The question is, is it a billing problem? It's deemed that it's a yes or no and given it a descriptive instructions.
So, the customer is reporting a billing problem, exactly the criteria that we wanted to assess against. The true criteria is the message reports a problem with charges, payments, invoices, refunds, or subscription billing. The false is the message does not report a billing problem. So, you can see that these are very descriptive criteria that Jev can therefore determine is a message or the state likely to be a yes or no against that criteria.
Claude has created this criteria based on the type-safe skill. So, it's come back with an answer then and the answer from Jev was effectively this one number, 0.99, where one means it was absolutely a billing problem and zero means it was not a billing problem at all. You can see that this is the response that we get directly from Jev. The input tokens were 324 tokens. The output tokens were 24. I also asked Claude to extract the time it took as well as the cost.
Jev's own share of that was 0.3 seconds and that's probably an overestimate based on the fact that Claude actually had to make the connection in the first place. And 324 input tokens is effectively 0.0000136 dollars. So, you can see how it could have done this for thousands of emails, and it barely scratched the surface on a couple of cents. You'd need 73,000 calls like this to spend a dollar. Extremely cheap, extremely fast.
And then the 24 output tokens, which was effectively just telling us the number, cost nothing to us. And by the way, if we jump back to what this skill actually does, the type safe skill, it doesn't actually run anything itself. It's basically a rule book that Claude can use or another agent can use with the three question types. So, it's able to determine which question type it should be in, how to write a question Jev won't misread.
So, it's about specificity of a question, when to batch questions together versus when to keep them separate, and we'll cover that later, and where to keep the thresholds, the confidence levels. So, if you're working from Claude, you never talk to Jev directly. You describe the job in plain English. Claude is going to read the skill. Claude's going to write the scripts to assess it against, and the script is going to send your rows to Jev, and Claude is going to read you the results back.
So, this is how you use them in conjunction. And alternative right is you could use something like Open Router, skip Claude code as a harness here, and directly use the model from Open Router, the Jev Lab model from Open Router. But that would effectively mean that actually you need to make up the questions, you need to make up the descriptions, you need to effectively structure that input to Jev there. So, you've seen now that the key works.
You've seen it's really quick. You can see we can scale this up. So, let's actually deconstruct how you would use this for yourself using an example of 50 support emails that come in overnight. So, I've got this CSV or Google Sheet here of 50 different support emails, when they were received, the ID of the email, when that customer has been a customer since, which plan type they're on, what invoice they last paid, the subject, the body.
So, everything that we'd need to pass into Jev. We're going to deconstruct on exactly how you would turn this into a workflow that sorts them so that the easy ones are answered, the risky ones are sent to a member of customer support, and that when they're queued, they're actually in priority order so they're ranked when somebody sits down in the morning to read these customer service messages. So, these emails are in a emails.csv that we made earlier, but equally you could if you're using a harness like Claude code, use the MCP to plug this directly into your email inbox and process this in the same way.
So, you can see we've got examples like "Hi, I've just seen two charges of $99. I signed up for a trial and the magic link never arrived." A feature request here, "Can you add perplexity to it?" Some customer complaints here like "Metrics haven't been shown." or "Metrics have gone down." Some questions about GDPR compliance. All of that is effectively data that we can go through and get Jev to assess extremely quickly against the set of questions.
Now, when deconstructing it, the most important thing before we write a single question is to decide what the code should do with each answer. So, the answers we're going to get back from Jev, the probabilities, the numbers, should all be a trigger for an action. So, every question that we make should be a trigger to a specific action. And we're going to talk about how to work out what questions to ask in the first place, and that's why we're going to go through and deconstruct this step-by-step.
And when we consider the different actions that we will take off the back of the results in this customer support case, we've got four different actions effectively. We're going to answer the email with Claude directly or draft it with Claude. We're going to put it in the engineering queue as something that's a bug that needs fixing. We're going to log it as a feature request for the future. Or we're going to escalate it or send it to a person.
So, that's then what I need to give to Claude, the job I want it to do, the four different actions or routes it can take, and one instruction, which is "Show me the plan before you run anything." And therefore, we're going to see effectively the structured plan that it's going to send to Jev, and we can go through and read before we do this on a bulk set of data piece by piece what the questions are, what set of data it's going to run on, etc.
So, let's talk about the prompt, and I've been quite specific here, but you can actually get Claude, like I did, to write a similar prompt for your exact use case. So, again, we've got "Use the TypeSafe skill." Very important. "I've got 50 support emails in emails.csv." I'm going to save it in my downloads. "Every morning I want them sorted into four piles, and these are our four actions, right? Reply with Claude, engineering queue, send to person, etc.
Each pile sorted with the angriest customer first. So, from that and using the type safe skill, it's going to understand the list of priorities and what questions we need to ask. And then it's going to understand that we need a scoring question to score it by, you know, anger, for example. And because I've got some idea of exactly what I want to understand for each, I've actually specified it here, but you can leave this open for Claude to actually determine the questions and criteria, too.
But for each email, I want to know what it's about. So, this might be a choice, billing, technical, feature request, account, or none of those. And it's really critical you put none of those in there because if it's not one of those, then Jev is effectively going to just try and pick the most likely outcome. Whereas if none of those is in there, it effectively gives it an out. I want to know if they're asking for money back or a refund.
I want to know if something is broken because therefore we want to send it as the action to put it in the engineering queue. I want to know on a scale, probably, how angry they are from calm, annoyed, or furious, so like one to three. And I want to know whether it's just something quite routine and standard, needs some judgment, or is unusual enough that a person should handle it. And then I've added some rules for what the code does or what the code should do.
Claude's code is going to create the code that's going to process Jev's answers. Anything about money should go to a person, never Claude. And I basically want a high confidence level on that one. If Jev isn't confident about the category or the case is unusual, send it to a person. If the product is broken, so they can't use it, it should go to engineering. Feature requests are getting logged. Everything else Claude drafts a reply using a a given reply guide, a separate brand resource, effectively.
And we're going to sort every pile by a priority score that's mostly anger, partly severity. So, it needs to score based on multiple questions. And Claude, working with the type safe skill, is going to understand that it needs to break these down into individual deterministic or more deterministic questions that Jev can then give us a probability score on and a confidence score on. And then we specified this because we want to do an extra check.
So, before any Claude draft is saved, check it with Jev in another request, a separate request, asking different questions against these questions. Does it actually answer the customer's question? And does it promise a refund or discount? So, we're doing an extra validation pass after Claude for the ones that are going to be written by Claude. And then we're only going to keep drafts that pass both. Write the plan in three sections, the state which is, you know, the input information, the questions with their criteria, and the routing with a threshold for every rule.
So, this is the confidence level that passes that given rule. Keep that in one file. Show me the plan before you run anything. And you could just ask to see the scripts that it's going to send to Type Safe or Jev, but this way we can copy this, paste it into our Claude code instance with Type Safe, and actually start running that. And what it's going to do is effectively come back with a plan, right? So, at this point we're less concerned with how quick it's going to run because what we're actually just trying to do is come up with a good plan with a good set of questions so that we can run it effectively using Jev.
Jev is only going to be really powerful for these bulk workflows and useful if you can ask good questions and Claude or an LLM is able to break down our requirements into those good questions. So, we've now got Claude's plan back in the questions.py file. And the first thing that we can see in Claude's plan is the state or what each email looks like when it reaches Jev. So, it's used the skill to determine what this state should look like.
Now, you can see that we've got ID, received at, plan, subject, body, all these different fields. So, what it said is actually we need one object per email rather than a flat string where each part gets a descriptive name so the relationship stay clear. So, it's going to include customer facts like the plan or the customer they've or when they've been a customer since because the founding customer on the top plan is a completely different scenario from a day-old trial and Jev will want to treat that differently.
So, it's basically mapped out the CSV columns here and given it some scripts. So, it's going to turn one CSV row into a state or input for a triage request. So, every single row is going to have information about the email which is going to be passed into Jev and information about the customer. So, it's going to be broken down into these specific sections. And that's because the skill told it to so that a question can point at a field by name and it's left everything that's irrelevant out of that like the previous thread history because the more irrelevant text it could potentially bring into the the decision-making process.
And that brings us to kind of the limitation or the hard limit of 64,000 tokens per request that you send to Jev, but you should never be near that. It's not designed like an LLM where you have a million tokens in the input. It's it's very different. These structured requests are, as they sound, very structured, to the point, and minimal. They ask very defined sets of questions. We've also got a state for the second pass, which is actually based on the subject and body of the email that's been drafted.
And we've also got a second state for that second pass that we're going to do when Claude drafts a reply. So, it's going to contain information about the email and the actual draft reply to see if we've answered the question, which was the second pass that we were going to do. The second thing that it's broken down here are the questions. So, we've got five questions per email in one single request. So, every single email in that CSV is going to get all five questions answered in one single request, and that's fine because every single question we add to that, we could have 10, 15 questions, a basically free.
So, we can continue to add these defined questions for free, and that's why we send them all at once so we can have this parallel processing of all of the different criteria. And in each response, we're going to get a score level and a confidence level back from Jev. And you can see it's built out these different triage questions. So, the first one we've got is a choice question. It says, "What is the support email primarily about?" If it's about billing, then the criteria is about charges, invoices, payments, subscription cost, but it's not for things like login or access problems.
So, we're being very specific what it is and what it isn't, and it's given some examples in there. Then we've got technical, feature request, account, and none. And like we said, in every request, we want to give it an out. We want to give it a none so that it can actually choose that rather than trying to just pigeonhole it into one of these categories. And then we've broken it down for things like asking for money back.
And again, we've got the criteria here for a null. If it's true, the email asks for money to be returned or credited. If it's false, the email makes no claim on money already paid. And it says, "Asking to cancel without mentioning a refund is false." So, they're not going to get that money back for that. We've broken down how badly broken and that is a score. So, the score gives the different criteria from effectively one to four.
Nothing is broken, it's a cosmetic damage, it's annoying, and they cannot use the product. Now, what's really important here is we're really specific with the criteria. So, what does annoying actually mean? It means a feature is degraded or fails sometimes and there is a workaround. So, because we specified that, Jeff can be more sure in its determination of whether it's annoying or not annoying. And it's important to remember that every single email gets put through the same criteria.
So, for example, a billing email will still go through and get asked these breakage severity questions. So, how badly is the product broken for this customer, judging by what the email describes? It will basically come back and say, "It's not at all." and therefore it will get ignored. But by sending all the questions at once, we can speed up, they're all parallel processed, we can basically just understand that it's nothing to do with breaking and therefore it's zero.
And the actions are not taken off the back of that score then for that email. And then for anger, how angry is the customer, we've got the score. So, how angry does this customer sound? We've got the score from calm annoyed to furious. Then we've also or it Claude has also broken out how much judgment it needs. How much human judgment does this answer need? Is it a standard case? Is it a judgment or is it unusual? And that is a choice there.
So, it's an option to choose any of those three. And the second part, we've got the verification questions around does it answer the question? Does it promise a refund or discount? And Claude is going to basically use the type safe skill to turn all of this information into a request that then gets sent to Jeff in the JSON structure that we saw earlier. But also inside here, inside the plan, we've got the routing rules.
So, is it going to get sent to a person, engineering team? Is it going to be put into a pile for Claude to actually go and draft? It's got the different confidence levels that we're measuring this against and you can come in and change these. And it defines those confidence levels. So, if it's below this category confidence on rule number two, then it's not going to be trusted as a result and actually it's going to be routed to a person.
And then we've actually got down here all the different routings in terms of actual scripts and Claude has gone ahead and written those scripts based on the type safe skill. You can see, like we said, there's no LLM determination here. It's purely a script. It will literally send all of the results from Jeb through these routes to actually put them into the right piles. And these are all if statements that are verifiable against the numbers that Jeb has returned as the output.
So, you've seen now how to write this using Claude, and that's going to be a lot quicker than writing it yourself. But, the harder bit actually is knowing which questions to ask in the first place about your process. And this is the bit that I couldn't find online. Nobody has written up. So, I decided to make my own checklist that you can grab for free in the description below that you can go through to try and work out what questions you need to ask of your process or how to come up with the requirements that you can input as that initial prompt.
So, this isn't even available in this skill that they provide for you. That's a very specific set of criteria of how to structure the requests to Jeb. But, here's the checklist that I've used for the different five question areas. And I've kind of pulled this together from the docs that I've read and the videos I've seen around the use cases of Jeb here. So, always think about it from what the software is going to do next, not from the data point of view.
So, every question that you ask Jeb exists because an action or it needs to trigger an action. An action depends on it. So, you're going to list the actions first. That's number one. So, ours, if you remember, were Claude replies, uh going to put it into an engineering queue, going to put it into a feature log, or send it to a person, route it to a person. And then for each action, you're going to write the trigger as a sentence.
And that sentence is going to form your question effectively. Let's use the example of send it to a person when it's about money. And that sentence is the question, and it also tells you the action, which tells you the type. So, if it's going to act or not act, it's a null or yes/no. If it's going to choose a different destination, then it's a choice. It's like a routing problem. Or if it's sorting, then it's a score.
Thirdly, we're going to write down what a person would look at before making that call. I.e., what is the input? What's the state? And in this case, it was an individual email or a list of emails. And then number four, we're going to add what else could be true that would change the routing. So, this These are your edge cases, even if they're super rare. And that's where bug severity came from. We're going to re-route depending on bug severity and we're going to actually then score that.
And then number five is deciding what a wrong answer is going to cost you so that you can set those confidence levels or those tolerances. So that's going to set the threshold for each action. And you saw that we had all of our different confidence thresholds. Now I've gone ahead and run this on all 50 emails in that CSV. It took 4.2 seconds for all 50 to be run using eight concurrent workers so it's done that in parallel.
It would have taken 31 seconds sequentially which is how you probably do it with an LLM. It would have taken 31 seconds sequentially so it's able to actually do those all at once. And the cost was 0.0026 for all 50 which is roughly equivalent to 19,000 emails equaling a dollar. So if we run this every morning for 50 emails it's going to cost us about a dollar a year. So extremely low. It's put it into this markdown file for us to easily digest the results and it's told us it's sorted by priority.
So it's basically said 70% anger, 30% severity in terms of the weighting it's given it on priority. You can see that 19 of them rooted to a person based on my anger and severity. And it gives us, you know, the subject, why it's been categorized as that. So this for example is somebody asking for their money back. On a scale of naught to three I think it was they are super angry and it's clearly something severe or it's been deemed severe by Jev.
And therefore our customer support person or customer support rep has a prioritized list of cases to go through every single morning and it took less than 10 seconds to work that out. We've then got a bunch of requests in the engineering queue that have been rooted there, some feature requests and 19 to actually just be replied to by Claude. Stuff that is no longer a task for our customer service rep. So now you've seen how you compare Jev with Claude, here are the rules you need to follow to get the most out of it.
These are the patterns taken directly from the docs. So number one is actually kind of counterintuitive but super interesting insight which is ask everything at once. We're going to ask all five questions like we did earlier in one request. So the problem this is going to solve is waiting for it and this is why it's so incredibly quick. So with a chat model you're going to ask one thing, wait, then decide what to ask next.
But, with Javelin, every question in a request runs in parallel, like we saw, and you only pay to read the text once. So, you ask everything up front, including questions that might not apply, like the bug severity for our billing scenario earlier, and then it ignores what you don't need. And this is called speculative fan out. So, for example, if it comes back as a billing query, the bug answers just get completely ignored in the processing.
And if it was a bug, you already have the severity, and you don't have to make another call to understand the severity. There's one request, multiple questions in that request, multiple types of request, null, choice, scored. Pattern two is gating decisions based on your confidence level for specific actions will help you build systems that are both reliable and safe. Now, I know that was quite wordy. An example that Typeface used directly in their docs is voice banking.
So, let's say under 0.6 confidence, so 60% confidence, nearly 50/50, it's going to get redirected to a person. So, check my balance at 0.6 is fine because the worst-case scenario is somebody hears their balance if it goes wrong. But, we need to gate different actions differently. So, approve the transfer needs to be a lot more confident, say 0.85, 0.95, and below that it asks you to confirm if that is the action you're wanting to take first.
And you can see this in the demo from the voice browser from the start. Scroll down just happens. It just scrolls down immediately. But, anything destructive, like closing a tab, is going to stop and ask you to confirm before it does so. And that's because it falls in below the confidence threshold that they've set. So, you have one confidence threshold per action set by what it costs you to actually be wrong on that action.
Now, pattern three is basically scoring the individual parts by breaking it down into into individual questions and weighing them or adding them together. And this rule or pattern exists because you're encouraged to break the judgment of these different things, independent dimensions, into separate questions that we can score each one separately and then combine them with weights so that we can then control the routing in the code.
And I'll give you an example to bring this to life, which is big, messy judgments. So, we could have a bunch of CVs, right? And we could ask, is this a good candidate? Score it 1 to 10. And you get one number from Jev that you can't question. So, it's hard to know what criteria is used to determine the outcome there. But what we might do instead is break that down into what good might look like. So, it might be for engineer, Python depth of knowledge, team leadership skills, system design, or how fast they pick up new things.
So, for a senior engineer, you could weight the Python skills and the design at 40% each as the threshold. But for a manager position, maybe leadership gets 40% as the threshold. Now, pattern four then in the final pattern is routing it to the best handler for that job. Or this is called intent routing. So, we're going to classify incoming requests and route each to the optimal handler, whether that's deterministic logic inside a script like we saw, a specialist LLM like Claude, or a human.
So, TypeSafe can sit in front of all these as a fast, cheap classifier that determines which handler to invoke, whether to send it to our human customer support rep or not. So, like where's my order is going to go to a database look up. That's not going to use AI. It's not going to require a human. But a product question is going to go to Claude who has access to the product docs. And then a complaint is going to get a second question of how complex is this complaint.
And then the messy ones are going to go completely to a person from the get-go. And then anything with low confidence is not going to hit the threshold and it's going to go to a person anyway. And when you bring all of those four patterns together, you get the smart home demo that TypeSafe have directly on their docs that uses all of this at once. So, this is about controlling your lights and your gadgets in your house using voice effectively or using typing commands or using an app.
So, you type in turn off all the lights in the house. One request is going to ask which kind of request is it, which rooms, which device, and what action needs to take place. And then if you type two commands in one sentence, Jev is going to spot that it's a compound request and actually hand it to an LLM to split that. It's then going to judge the two halves there. So, it understands that it needs to actually split the request into multiple questions there.
And then if you ask just a general question that's unrelated, it knows to pass it to an LLM to answer that. And that is effectively pattern four. And the whole point in that is that they put Jev next to the LLM, and you basically Jev works so quick on these deterministic actions where you've got questions, threshold, confidence, etc. That you don't even notice he's there. It can make judgment calls on things so quickly that decide the next most probable step to take depending on score and confidence against your criteria.
So, if you can't tell, I am super excited about this because when you pair it up with Claw code or GPT-Astra, you've got one hell of a combination. Jev can do a lot of those deterministic or more deterministic categorizations in rapid time and for record low cost. Whilst Claw can do the chat back and forth, the reasoning and the judgment where extra thinking or looser requirements, for example, are required. Now, if you want to see a video on applying this combination to a domain like GEO or browser use, then drop a comment below and see you in the next video.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.