Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Dave Ebbelaar · @daveebbelaar
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in Dave Ebbelaar's most watched videos.
Most replayed moment at 15:10
12.3x that video's typical replay level
of things that you would like to have there then give a detailed report of what we can improve the llm will do that you now have a list of feedback points and we feed that back to the llm here's the original output here's the feedback now process all of that all right so
Said at 15:04
Most replayed moment at 14:01
2.4x that video's typical replay level
what do we need oh the user wants latitude and longitude oh I have to know the latitude and longitude of Paris let's give that back all right so now what we can do we now have the parameters that we need to plug into this function because we have the
Said at 13:55
The graph counts replays. It does not show where viewers stopped watching.
Words
3,235
Runtime
16:50
Speaking pace
192wpm
Reading time
13min
192 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
In this video, you are going to learn everything you need to know as a developer about Type Safe's new Jeff model. So, let's dive into things. This has been making some waves and we finally have something happening in the AI industry that is super relevant for developers building AI applications. Because for the past year or so, all the focus has been on the coding agents, the better models. And of course, while those are amazing, when it comes to building LLM based systems, agentic systems, whatever you want to call it, we have
96 words, the words spoken in the first 30 seconds at 192 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 161 |
| Average words per sentence | 20.1 |
| Longest sentence | 90 words |
| Questions asked | 28 |
| Sentences containing a number | 12 |
Most used terms
Filler phrases
56 in total: like 26 · right? 19 · actually 5 · you know 4 · I mean 1 · kind of 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
In this video, you are going to learn everything you need to know as a developer about Type Safe's new Jeff model. So, let's dive into things. This has been making some waves and we finally have something happening in the AI industry that is super relevant for developers building AI applications. Because for the past year or so, all the focus has been on the coding agents, the better models. And of course, while those are amazing, when it comes to building LLM based systems, agentic systems, whatever you want to call it, we have pretty much been working with the same tools since function calling.
For the past 2 years, while the models get better, we haven't had a new paradigm or a new model to really play with. Jeff is the first of its kind or system one is the model they call it is the first of its kind in this new category. So, I am going to walk you through a couple of practical examples in this repository that I have created for you. We're going to walk through some code examples and not just cover like the blog posts and the highlights that everyone else is covering, right?
But show you in a code base how it works and how you can build AI applications around this. So, what should you know about this model? Well, the first thing is Jeff system one, it's a classification model. That is the starting point. So, you can ask it a question and it will classify your inputs and outputs either in a yes or no, or it will select from a series of options that you give it, or it can also give a distribution really of the confidence scores that it thinks the criteria has.
So, this is super useful when you are building applications with AI. Because as you know, if you've been building with AI, with LLMs, what we've pretty much been doing in the past is we had these language models that could given, let's say, a user question, could give these like outputs in the form of text. So, this would be the response, right? So, early in 2023, something around that time, we figured that it would actually be very useful if we could do function calling using these LLMs and use structured output in the form of let's say JSON objects.
Because when you are building programs and you ask it something you can get a JSON object back, that is when we can make decisions. That is what that when we can apply if else, we can route, we can do all kinds of cool things as programmers in order to do this. But the most important thing to understand about this is that LLMs were not trained or optimized to do this. It's kind of like a second property of the large language models.
And what the team of Type Safe did is essentially they flipped that paradigm around and they said, "Look, we need something else. We need a way for these models to always give JSON output and to work with that." So, that's what they started building and essentially they took a large language model pre-trained language language model as the input, but they have an additional like other training layer at the end of it, on top of it, to make the model only be able to work like this.
So, now with that interaction out of the way, let's make this a little bit more practical because you just need to see some code examples and play with it in order to get a feel for this. So, we are currently using the Python SDK and the link to these exact code files, it will be in the description. It's in my AI cookbook for you to follow along. So, what I'm going to do is I'm just going to walk through this and show you some of the examples and the data types and inputs and outputs that you can expect because I think this is all you need to know at this point, inputs, outputs, and the choices, and from there you can decide how to implement this into your applications and just let your AI agents build all of this.
So, very simple, we got the Type Safe SDK, we have a client, we load our API key. If you go to the console, you can get the API key there. Currently, you need to sign up for the waitlist, but pretty much from everyone around me, you'll get access in a couple of hours, and you can play along with this. So, we have this client, and similar to how we work with chat GPT or Anthropic, you know, we understand how this works.
We can call client, and then we call dot system one. Then we have our inputs, and then we can get a response. So, let's just first run through this and see what we got there, right? So, if I now print this, we get a response back, and that says billing. So, let's look at what's happening over here. So, first of all, you can see there is a state over here. So, I was charged twice. Please run refund the duplicate. So, this is the starting state that we want to know something about, right?
And then we can ask a question. And we'll get into the data models and the options later in these other files, but for now, we just give it a question, and that is a choice. And that question is which team should handle this support ticket. And then we give it three options. It can be billing, technical, or other. So, we run this, and then you see billing. Okay. So, up until this point, if you've been working with LLMs and structured output, like this is it, right?
This is structured output. It's very simple. There's nothing new that we can do here that we couldn't do before. So, then what's interesting about Jeff? Why is this all the hype? Well, first of all, because it is this new modeling paradigm, we get two extra properties. First of all, it's fast, and second, it's very cheap. Now, that is of course from an application and a programmer perspective, cheap and fast, that's what we want, right?
That's that's great. So, that is the whole philosophy and the promise of Jeff that we use AI inside applications, and we can make decisions, smart if else statements, if you will, and we can now do that way faster and most importantly, way cheaper than we could with some of the other models, even if we use something like Claude Haiku, for example. It's about 20 times 25 times cheaper than Haiku and also faster. So, that's the introduction to Jeff.
This is how you work with the SDK and how we can classify an input. So, let's now look at the different types of questions that you can ask, right? So, if I go over to the documentation over here, this is I think the most important thing to understand about this. You have choices, you have scores, and you have nulls. So, let's start with the choices. This is what I just showed in the most simple simple example, really the most easiest probably to understand.
And that is you give it an input state, I was charged twice for this order, please refund the duplicate, and you give it a a list of criteria, and those are the categories that you can choose from. And your instructions is the input prompt. So, rather than thinking in system messages and system prompts and user messages and then structured output data models, the whole API and also the models are designed just for classification.
So, I would say this is a like third property next to like the speed and and the cost really that make this a more interesting model. Also from like a programmatic perspective, the API and the model itself forces you to structure information simply in such a way where you don't really have to think about large system prompts and how to structure all of your information and in what format because the API is is very declarative in a way saying like, "Hey, here's here's what you need to input and here's what I can work with." And you don't have that much variety in there.
So, if I run through this example over here, you can see that we import the choice from the type safe SDK. We create our input state, which is in this case just a string. Then we create the question and we set that equal to the choice with the instructions and the criteria. And then here we can actually call client.system1 again. We put the state in here and also the question which we can also give a name with another string and we can run that and we get again we get billing but now we can also see those probabilities and we can even get a confidence score.
So, that's the choice. The second one is a score. This is similar to a choice but rather than providing it with a dictionary of categories and a description, we just give it a list. We just give it a list of inputs criteria, I should say, where it can choose from and rather than picking one, it will give a score for all of these different categories in it. So, here we input the score. We can put that into a question and then when we run this and we ask about these different criteria, you can see the probability distribution over these categories.
So, you can see that the confidence score right now is one for frustrated but civil. So, this is an easy one for the model to classify but if we start to have more ambiguous messages or we have more categories in here, then you will see that the probability distribution is going to shift a little bit. Let's now go into the null which I'm not really sure why they took this name null. It might be coming from like a Bernoulli distribution but this is pretty much a Boolean.
Yes or no, true or false and it will also give you a confidence score on this where the score that it will give you is the probability of it being yes. So, let me show you what that looks like. So, again we have the example I was charged twice for an order, please refund the duplicate. So, this is the ticket so we're going to use that as a state and then we have a null and the question is does the customer explicitly ask for money to be returned?
So, this is a this is a yes or no question. Let's put that in here and let's run this and look at the response. So, here you can see we get the answer. So, the refund requested is a null answer type, and the null is 0.98. Meaning that is the 0.98 0.98 probability that the model classifies this as a yes. And now, based on that, we can perform all kinds of logic if else statements, right? So, that is really the starting point for you to start thinking about this model.
It is a classification model, and depending on whether you have different choices that you wanted to select from, a yes or no, or certain options where it needs to calculate a probability distribution over, that is what we can do with Jeff and with the system one model. Now, let's look at a quick speed comparison to see how fast this model really is right now, and let's compare it to some of the Entropic models. So, we do Haiku, we do Opus 5, and Fable 5.1.
So, if I start running this here on the side, let's see. You can see the runs over here, and I am based in Europe. You can now get the fastest speed through the results really if you are on the in the US, because here you can see on average it will take about like 500 600 milliseconds for a request to complete. And if we compare that to Haiku, which is also a great model to perform classification steps like this, you can see it's not that much faster, right?
So, I believe like if you go closer to the source and and and where where you actually run this model, I think it can be way faster. But, what is more interesting right now as a developer is probably the cost, right? So, if you go over to the pricing for Jeff, you can find the official model pricing on their website to figure out what this currently looks like. But, to give an example, if you perform one 100,000 requests with about like 500 input tokens, running Jeff that will cost you about $2, whereas with Haiku, which is already cheap model, that will cost you 50 bucks, right?
So, that is a 25X on cost. And the interesting thing about this is that this model is the the first of its kind, right? So, this price will go down significantly. And as developers building AI applications, I think this is one of the most interesting aspects of it that we now have new models specifically optimized to put them into your applications and to make fast and cheap decisions. So, what can you do with Jeff's system one and what you should use it for?
Well, on their documentation, they also explain a couple of patterns that you can use along with like the benefits, right? But, if you've been building with LLMs and if you use structured output in the past, then this is a very natural extension of the work you've already been doing and you can probably identify where this makes sense, right? One of the cool things about the Jeff model, the system one model as well, is that you can also combine questions.
So, the API supports this. So, let me show you what this looks like. So, here you can see we have a choice, we have a score, and we have a null. And then we start with a ticket, which is the input state. Then we just create a dictionary of all the types of questions. And then we can essentially make one API call and we have all of that information here available. So, it will decide, look, what is the category? What is the frustration level?
And is there a refund request? And now, as you start to build bigger, more complex systems and you start to combine multiple choices, scores, whatever type of classification what you want to do. Whenever you stack more together, I think that's where a system like this is going to really outweigh the current options that we have available as developers when using large language models. So, not only is the API around it easier, so you will have less ambiguous code, right?
Less like fake system prompts that people need to change and oversee. This is much more programmatic and because it's now faster and also getting cheaper, it will also start to make sense to just put more and more of those choices into one system and just for the sake of like being very certain, create all these conditions and branch out your applications or your workflows in a way where you can get really granular on if this happens combined with five other things that also needs to be true, that is when we do that is when we do this.
And if you've been following my channel, this is pretty much what I've been advocating ever since we've been building with language models that the best way to build the most reliable applications is still to have these workflow patterns, to branch off, to have these routers, these if statements. And these patterns are just combinations of combining the choices, the nulls, and the scores in a way to get to the final decision, the leaf node really of your workflow to then make a decision, perform an action, call a function, whatever makes sense for your application.
So, Type Safe's system one model, what's the verdict on it and should you use it? Well, I think this is one of those things where in a couple of years we're going to look back at this and see like, "Yeah, that was a pivotal moment in AI." But, for me it's very early, right? This is the first company, the first model of its type. It's a new company. For me, that's always tricky because I build AI solutions for clients and data privacy is always a big concern, right?
So, I know they have their data policy here. So, I haven't really gone through this yet, but even if the policy says everything is fine, explaining to a client here in the Netherlands or Europe, for example, that you now also need to share your data with a new US company is always going to be a little bit trickier, right? So, we all have all of that covered in Microsoft Azure. We still use the models from OpenAI, you know, but that has been established now.
So, I think over the next 3 months all of the big labs will have a similar version of what this has to offer. OpenAI, Anthropic, 3 months max, I think they have something similar. I think that's a good thing. So, that will probably drive down the cost even further, probably make it faster. Eventually, I hope that open source will catch up and that we can get smaller models that we can actually self-host. I mean, that would be so amazing, right?
Where we can use the big models for reasoning and for the big task, but then have simple small open source models that perform this type of classification with very high precision. So, I think this is going to kick off a new trend. I think this is going to be amazing for computer use, for browser use, which is going to be really big over the next year and years to come. So, I think this is really exciting and I'm really curious to see what's going to happen with these developments over the next months.
So, that's it for this video. Now, if you want to know what types of projects and solutions I build for my clients, you can check out the link in the description to check out my proposal system. It's completely for free. You'll get the template that I use to create proposals with one anonymized client proposal that we actually delivered. So, that gives you a chance to see what it's like to work as a freelancer to build with AI, which is one of the greatest things you can do as a developer right now next to your job.
And with that, I am going to wrap up this video and show one video over here that I think you will like next.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.