Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Venelin Valkov · @venelin_valkov
Words
3,092
Runtime
19:09
Speaking pace
161wpm
Reading time
13min
161 words per minute, between the 160 25th percentile and the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Is it time for birdlike models to make a comeback? And are general purpose classifiers going to be the next big thing into the AI world? Well, for the first one, it seems like this is the case. For the second, I'm not really sure yet, but I really hope that models such as this one that we're going to be taking a look at today, Jeff is going to be having open weight and open-source alternatives that are going to
81 words, the words spoken in the first 30 seconds at 161 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 178 |
| Average words per sentence | 17.4 |
| Longest sentence | 142 words |
| Questions asked | 20 |
| Sentences containing a number | 24 |
Most used terms
Filler phrases
31 in total: uh 15 · like 11 · um 2 · I mean 1 · actually 1 · basically 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
Is it time for birdlike models to make a comeback? And are general purpose classifiers going to be the next big thing into the AI world? Well, for the first one, it seems like this is the case. For the second, I'm not really sure yet, but I really hope that models such as this one that we're going to be taking a look at today, Jeff is going to be having open weight and open-source alternatives that are going to be matching its performance. and even get better.
But for today, let's try Jeff. Let's get started. This is the announcement post from Typesafe AI introducing the system one models that is the general category of the models and in particular the first one of those called Jeff. They were working in stealth for about 2 years. So I'm not that familiar with the company or their research if they have published anything at all. and they're saying that they have built a complete new stack, a new model architecture which is very interesting parallel sample.
So every time when I hear a sample I'm thinking about real and useful probabilities uh compared to what we are going to get from so you're going to be able to calibrate those within the outputs and the predictions and a new training method we called reinforcement warning for calibrated decisions. I haven't been looking into any papers or any research for this type of reinforcement warning. Unfortunately from what I have seen so far the guys from type safe I haven't published any research on what reinforcement warning for calibrated decision is and unfortunately again the Jeff models are not open and open source.
The first public model of the system one models is Jeff available today in early access and I was given access within 6 hours of applying to their web application. Jeff achieves similar in levels of intelligence on system one tasks compared to existing. So this is a bit let's say wowsy definition of what the performance of Jeff is while being two orders of magnitude faster and more efficient. Okay, we're going to check that one out.
While Jeff gives up string generation, it's optimized for structured output and cannot hallucinate. Yet, this is truly the case since this is more of a general classifier if you will compared to an OM. Think of Jeff as a frontier intelligence function call and structured state in type proistic decisions out. This is the comparison between the Jeff model and the existing OM. Note that here the training methods are vastly different.
The existing ALMs are optimized for human preference and trained with reinforcement learning with human feedback and other approaches of course. Then for reinforcement learning for calibrated decisions, their new approach to training is used for the Jeff models. Again, not really sure why what this means and how it is being implemented. And here they're saying that they're implemented and optimized for the human preferences.
That is the case with the MS while Jeff is for calibrated decisions answers with epistemic honest probabilities on system one tasks. Okay. Yeah, this is a general classifier again and the inputs unstructured data for the M text with an emphasis on sequential messages. Uh here of course you can pass a wide variety of inputs, text, images, videos and other media if you will audio messages as well in most while the Jeff models is working only currently at least with unstructured data with an emphasis on structured program state.
Okay, this is uh good and the outputs of course from a mostly we're getting generated text which are generated one token at a time and from the GIF models we are essentially getting type safe structured values that are specifying here those are the outputs provided for the input state that is going to be given as a classification for the different tasks that the model has support and the sampling here is sequential since we're generating one token at a time in most while it is parallel of course again what a quasifier is doing under the hood and then here they have a very good comparison on input price and output price of course are quite expensive but they can be pretty powerful while the Jeff models are extremely cheap talking about 0.42 for two cents per million tokens.
I guess that this is some way of play with the pricing right here. Maybe they could have done this even cheaper. Not really sure about that. And the output or the predictions of the models are essentially free while we are paying quite a bit of money for AOM's outputs. And since this is a classifier, uh we're getting much faster responses. I would assume that this model is within the range of 500 million parameters to let's say 7 or 8 billion.
And if we assume they are not subsidizing their predictions, I would guess that this range is in comfortable enough for them in order to provide very good inference and inference time. They have some workflow evolations with front tier models and models such as GP 5.6, SOA, and Opus 5 and Sonet 5. I'm not really going to comment on these comparisons since this seems like a comparison between oranges and apples. And in my case, I'm not really sure that this is a fair comparison.
I would more likely look into independent comparisons than with this model. And of course, the tasks that LMS are doing are vastly different compared to what you might do with GF type of models. They also have these comparisons between pretty much every possible frontier on structured output error rate and to call error rate. And you can see that Jeff is at 0% that is doing perfectly on both. And this is expected and something that looks like a marketing trick that it is going to be playing around with your knowledge instead of looking at what these types of models are capable of.
Of course, Jeff models since they are unified or general classifiers, they're going to be having zero error rate on tool calling and structured output. When you sign up for Typesafe AI account, you're going to be given five bucks as credits and those credits with the Jeff usage is going to be giving you a plenty of play around with the model. So, go and do that. I was given in about six or so hours access to the model.
So you can do that same thing and they have a pretty nice Python SDK. It is called Typesafe SDK. And here they have a pretty good example of how you can initialize it. Type safe client. Then you're going to specify the system one client. And you're going to specify the state. Here they have an example with hi I've been trying to connect to my Strap account for 3 days and keep failing. I'm losing sales. Please help ASAP.
So here you have the model answer a couple of choices for example what department this should be handled with billing technical sales and u most of those are going to have some description. So basically all of those are defined by you the user and they have frustration how frustrated the customer appears and they have uh no which is yes or no the message conveys urgency or time sensitivities and here you can see the frustration score which is 1.035.
So in this case the result is frustrated but CU the scores are coming from 0 to two. So this is the calibration range for that and this is yes or no. So essentially this means it is urgent. If you want to become a better AI engineer, go and subscribe to my AI engineering academy on mexpert.io. There you're going to learn how you can build arax chatbots and agentic systems and how you can deploy them into production. So if you want to become a better AI engineer today, go and subscribe to M Expert Pro.
Thank you. I'm in my wok VS code and here I have a notebook that is uploaded to the GitHub repository that is going to be within the description of this video and I have awarded the Jeff model. I have it here and I'm also specifying the pricing which again is 0.42 cents per million tokens that is for the input while the output is essentially free. And here I have a type- safe client. I am initializing with a crew travels through wormhole.
The visuals were stunning. Despite that, I would recommend watching it. Here I have the state, which is essentially a movie review that I'm giving in with questions. Does the reviewer recommend watching the movie? So this is a yes or no question. And I have choices for the genre. Which genre best fits the movie described in the review? Note that these all these are defined again by us. And the input context window for the Jeff model is about 32 or 64k tokens.
Uh I mean I have found two different texts on that and not really sure about that but we are well within that range. And we have some enthusiasm. How enthusiastic is the reviewer? Again this is a score with criteria and three different options. Once you run all of this you're going to see the result which is does the reviewer recommend the movie? Yes, with the probability of 98% and here we're getting this calibrated probabilities that we're not getting from a what genre is the movie and the one selected by the model is sci-fi with confidence of one.
And here you can see the different scorings and sci-fi is 100% essentially. So this is a pretty good classifier at least on this review. How enthusiastic is the reviewer? Weighted score one. So out of these we are getting positive with reservations and this is again something I would highly agree with. We are getting the input tokens and the output tokens and we are seeing the uh overall output of that and let me show you the usage of this.
So this is like very very minimal price for this particular input. You can imagine that this is going to get quite uh cheap even when you're using into real world applications. So we have the choice versus the score versus no. And the choice is which of those alternatives fit? What does this sit on described orders level? So give me a score between zero and two and is this statement true for the no? So what is the probability of true?
The next example is a bit more dynamic if you will. I have a ticket for issue an online store with a stripe integration. Hi team checkout returns HTTP 500 for every customer after today's deployment. Nobody can pay and our launch starts in 20 minutes. Okay, so this is essentially the ticket. We have a triage of questions. Which team should handle it? Billing, engineering, sales, other frustration. How frustrated is the writer stone?
And is it urgent? Yes or no? So it is pretty similar to the quick start example here. And here you can see that we are getting a 0.6 seconds in total time in order to get the output of this one. And we are getting again very very cheap output for this one. And the results I'm going to show you uh those in a second. So for the choice of competing teams or which should be completing this is the engineering team with very high score confidence of one.
You can see here again very nicely calibrated probability for this one. What is the frustration on 0 to2? So here we're getting a very calm response. So um there is no frustration or no angriness. And is it urgent? Uh you can see that we are about 0.98 or 98% yes it is urgent. For the next example I'm going to be doing something like semantic ranking using the Jeff system for a fictional rack application. Have a query.
How can rack retrieval handle both exact product ids and paraphrased questions? So here we're seeing a couple of documents with hybrid design. We have hybrid retrieval architecture. This is essentially the architecture between semantic search and something like BM25 dense notebook install sentence transformer. So this is much more detailed into the implementation on how to do it. BM25 recipe again pretty highly detailed implementation on how to do it.
And then we have something completely unrelated. PyTorch image classifier tutorial. We have a couple of ranking rubrics. The document does not address the retrieval problem in the query. And for the implementation, we have again four different ranking rubrics. For conceptual discussion only, names components provides or explicitly links a runnable example with setup and simple inputs. So these are the rubrics which we are going to be evaluating the implementation.
I am preparing pretty much all these for the different uh options that we have within the documents. How much concrete implementation guidance does uh the particular document is providing provide or explicitly describe? Which implementation detail regardless of relevance? And then for the relevance, how well does document address query which is going to be the query that we are passing in as input. And I have a couple of examples here for hybrid design implementation and relevance.
So you can check the concrete data that is going to be passed to the model. And I'm asking Jeff with the search query, the documents and the questions that we have provided right now. So let me show you the actual result. And here you can see that for the hybrid design we have relevance to the query of one and implementation detail or how well the implementation is described is a very low score. So for example you can see um maybe the dense notebook dense retrieval notebook relevance is uh relatively high but not as close to one and then the implementation is very very specific the pytorch image classifier while it is very specific into the implementation the relevance to the query is pretty much zero so it shouldn't be taken into account and we have relevance confidence and implementation confidence and you can see that using this approach which might be something that is worth doing into practice and maybe it can even replace your ranker within your ARC systems.
And this is the combination of relevance and implementation. While you can see that for relevance of at least 80 or 0.8 we are getting a hybrid retrieval architecture to be number one. But if I decrease the relevance for example to 40, you can see that the runnable density retrieval notebook now the implementation contribution is much higher and we are seeing at first place that this should be the document that we're going to be requiring in order to answer the rack question.
Jeff can't work with images. I have devised another way to try and still work with some form of image data. I'm going to be using OCR in order to transcribe this particular receipt and I'm going to be asking to classify the wines one by one and tell me which of those is going to contain a total amount and these are the ones that were extracted by Tesseract from the open source OCR library and you can see that the extractions are definitely not perfect.
We have quite a bit of uh issues here. For example, you here we have ordo should have been London or something like that. Yeah, you can see that definitely this extraction wasn't perfect. Then I have essentially parsed the wines and have been looking for some uh digits or decimals and we have the values that were being found within the lines. Then I was essentially giving those wines to the Jeff model itself and it has been like quite cheap again.
The wall time was pretty minimal and it has selected candidate amount for which should be this one 9619 and you can see that this actually is the total amount at least on this receipt. So this is pretty encouraging I would say. Let me show you the example here. Jeff selected OCR span 9619 and I have essentially given a bounding box around it. So at least on this example it was pretty accurate on this one and this is a way to get around of not giving Jeff images but in this case we were working with OCR text.
I already have some open weight models that are trying to do the same thing that Jeff is doing. And here they're even using system one calibrated decisions with RL CD and yeah strictly proper scoring rules. So yeah, they have even some evolutions against Jeff and the interface that I'm seeing here agent predict is pretty similar to what you might get from a Jeff API call and yeah you can try these open models on your own and they in some case might be something that you're looking for and honestly they have pretty much exactly the same interface which is great and yeah we can be working with more bird or modern burn fight tunes like this and not be needing anything like Jeff that is going to be closed source.
So this is it for this video. We took a look at Jeff. In my opinion, this is a very interesting model that is going to pave the way for openweight models that are going to deliver this similar or even better performance compared to these quote source alternatives. And I assume that this is a relatively small model compared to most or frontier and the open- source world is going to be catching up very quickly with those types of models.
And hopefully those types of models are going to be more relevant into the real world since they're going to be having quite a bit of integration within the real world systems. Thank you for watching guys. Please like, share, and subscribe. Also join the Discord channel that I'm going to link down into the description of this video. And if you want to become a better AI engineer, go and subscribe to ML Expert Pro. Thank you and I see you in the next one.
Bye.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.