![Вот на что способен Jev. Три задачи для бизнеса [Гайд на Jev]: video thumbnail](https://i.ytimg.com/vi_webp/TTYefEu8sxw/maxresdefault.webp)
Вот на что способен Jev. Три задачи для бизнеса [Гайд на Jev] transcript
HinkoK · @hinkok_
Words
2,400
Runtime
14:46
Speaking pace
163wpm
Reading time
10min
163 words per minute, between the 160 25th percentile and the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
A week ago, the person behind the creation of ChatGPT developed a new and completely unique model in the AI world . It is called Jeff. And it really blew up the internet because it is hundreds of times faster than other language models and hundreds of times cheaper. And most importantly, it does not hallucinate at all. Well, at least that’s what they claim. But the problem is that 90%of people don't understand what it is at all, and they confuse
82 words, the words spoken in the first 30 seconds at 163 words per minute.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 203 |
| Average words per sentence | 11.8 |
| Longest sentence | 34 words |
| Questions asked | 9 |
| Sentences containing a number | 52 |
Most used terms
- jeff31
- model17
- requests15
- sonnet15
- request12
- haiku10
- task10
- questions9
- skill9
- answer8
- simple8
- works8
Filler phrases
13 in total: like 5 · actually 4 · literally 2 · kind of 1 · you know 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
Transcript
A week ago, the person behind the creation of ChatGPT developed a new and completely unique model in the AI world . It is called Jeff. And it really blew up the internet because it is hundreds of times faster than other language models and hundreds of times cheaper. And most importantly, it does not hallucinate at all. Well, at least that’s what they claim. But the problem is that 90%of people don't understand what it is at all, and they confuse it with regular ChatGPT or Claude.
But it is not a language model in the traditional sense at all. And it works in a completely different way. It is a decision-making model. In this video, I will break down in detail what Jeff actually is and how to use it. I will show three ways to apply it to real-world tasks. And most importantly, I will put together a ready-made skill for you that you can take completely for free. Well, before starting the video, I want to ask you to hit the like button.
It really helps me with promotion. Thank you. On September 15th , the company Typeface AI rolled out a brand-new model called Jeff. And it immediately spread across social media because it is not a regular language model, but something else entirely. And now I will explain what it is. It is not just a model that writes text or creates code. It actually cannot do that. It communicates in probabilities. Here is the best example I found on Twitter.
And after this, you will immediately understand how Jeff works. So, look at how it works. How does a normal LLM work? You send it a request, and it prepares an answer word by word. Sometimes this takes, well, a long time . In this case, it took 8.5 seconds. And this is how Jeff works. You send it the same request, and instead of writing word by word, it writes almost instantly in 0.1 seconds, giving you three probabilities.
The first is that this invoice is a scam, and the probability of that is 7%. The second is that this invoice is completely clean and can be paid. There is an 88% chance for that. And the third is that this invoice needs additional verification. There is a 5%probability for that. So it gives an answer in probabilities, and it does it hundreds of times faster, because the previous language model took 8.5 seconds to respond.
Jeff responded in 0.1. Due to the decision-making speed of this model , it has a multitude of applications. For example, it can play computer games in real-time and make various decisions . Or here is another option where you can use Jeff, and it just blows your mind. One enthusiast cobbled together an app where he made Jeff a full-fledged autopilot for a car, and it makes all the decisions on the road. it stops at a red light or continues driving if it is allowed to go and it isn't violating any traffic rules.
For example, it stops at a stop sign because that is the correct thing to do . In general, there are simply hundreds of ways to use this model. But the coolest thing is that it is extremely cheap. And now we will talk about the price. To help you understand how cheap it is, 1 million Jeff tokens cost 4 cents. Furthermore, 1 million output tokens are completely free because it spends so little on output that the creators decided not to charge for it at all.
However, the context of this model, unfortunately, is still only 32,000 tokens for now. And it can already be used either via OpenRouter or by registering on the official website. It was trained using a method called RLCD, which is tailored for calibrated decision-making rather than writing text. Essentially, it is an instant calculation instead of writing text word by word. Well, I think we have covered the theory, and now let me show you how it works in practice.
Moving on to the first case. And we will start with what I think is the most understandable task. Many businesses have it. It is incoming requests. For an example, I took a tech repair service where clients write from Telegram, email, and WhatsApp. And this business has accumulated 300 requests that need to be sorted into different categories. In this file, you can see the 300 requests that have already been exported.
These are exactly the ones we will be working with. I have prepared this website where it will be visually clear how Jeff handles the work. By the way, all materials, links, prompts, and skills from this video will be in my Telegram bot via the first link in the description. Go there, grab them, and test them yourself. So, if I click " random request from file," a random request appears from the 300 I showed you. For example, "refund my money for the technician's visit.""He looked at it for 5 minutes, said 'I don't fix this,' and took 20 dollars." Okay.
And I also have four questions that I ask Jeff so that he classifies this request based on these four points. Is this request spam? Well, most likely, no. Second is what category this request belongs to. For instance, a request about price, a refund, or other categories. Third, how urgent is this request? And fourth, is a human needed here? Well, let me finally press this button, and you will see how Jeff works in reality.
And it works very quickly. Look, in literally less than a second, in 670 milliseconds, he provided an answer to all four questions. First, 3% chance that this is spam. Meaning he replied that it is not spam. Next, in the category, he selected that it is a complaint. And it truly is a complaint. Also, there are options for purchase, support, price, and partnership. Next, urgency: respond today. And this really is the kind of complaint that is best answered today.
And it requires a human 70%—meaning he replied that yes, a human is needed here. And for this specific request, I think, yes, it's better for a human to answer than just an AI agent, because the situation isn't exactly standard. Besides, he spent 1,100 tokens on this, which cost literally nothing. But for this request , he only answered four questions. What happens if we add 19 more questions here? How much longer will it take him to answer?
Let’s check. I’m pressing the "Add 19 questions" button. And he handled it even faster, by a hundred or so milliseconds. True, because there were more questions, he spent slightly more tokens. But that was just one request. What if we give Jeff all 300 requests at once? How long will it take Jeff to process 300 requests? Let’s check. I hit the button, and the run begins. And here he is processing all the requests at a very high speed.
Okay, it took him 11.5 seconds. And he spent 1 cent on it. This is just mind-blowing, considering he processed all 300 requests. Now let’s check how well and how fast Sonnet handles this. I press the button to run Sonnet 3.5, and it begins. 11- plus seconds have passed, and he only processed 20 requests by the time Jeff finished all 300. And he has already spent more on this. Anyway, let's wait, and I will show you the final result.
And finally, after more than 2 minutes, Sonnet finished. And he spent $ 1.18 on it, which is more than 100 times more expensive than what Jeff did. He was also tens of times slower. Well, in my opinion, Jeff just crushes this and other language models here. But let's see how well Jeff and Sonnet handled this task, because speed and price are one thing, but accuracy of the answers is something else entirely. Jeff handled spam with 96%accuracy, meaning he got 4%wrong.
Sonnet was correct in 100%of cases. And here, Sonnet wins. Next, in category distribution, Jeff handled it at 96%. Sonnet at 94%. So Jeff actually did a little bit better here. Next is urgency. Haiku 64%, 65% —more or less the same, Sonnet is one percent better. And in the "human needed" category, Haiku answered correctly 89%of the time, Sonnet 93%. So, Sonnet is 4%better here as well. There is a difference of a couple of percent.
It's true, Sonnet is slightly better, but is it worth that much money —100 times more—and that much time? Personally, my answer is no, it's not. And by the way, this demonstration was real. I've logged into OpenRouter, and you can see that Sonnet spent over a dollar. Now, if you want to learn how to AI-code but don't know where to start, grab my free guide. In it, I first break down the tools themselves, Plot Code and Codeux, and then I build a real project from scratch and show the whole process in practice.
The guide is in the Telegram bot at the first link in the description. And now we move on to the second use case for Haiku. And I really like it because it saves us real money. In it, we will build a model router where Haiku will be responsible for choosing which model we should use for a specific task. Because if we ask a very strong model, like Opus, some simple questions, we'll just be burning through our money for no reason, since that same question could be asked to a weaker model.
And that is exactly what Haiku will do. It will distribute our tasks to the models . It will give the weak model easy tasks, and the powerful model it will give hard tasks. This is what this router looks like. And in reality, it's very simple. There are only three models here. That's Haiku, Sonnet, and Opus. But if you want me to build a real one that distributes tasks between OpenAI models, DeepSeek, and others, then write about it in the comments.
And there is a calculator here. If we make 1,000 requests a day and the share of simple requests is 60%, but we do all of this on Opus 3.5, then we will spend 420 dollars a month. If we do it through a router, the savings will be almost twofold. Look, I'm giving a simple prompt first. Calculate a 15% discount on 84 dollars. And theoretically, it should go to Haiku 3.5, because, well, that's too simple a task. That's exactly what happened.
Haiku identified the task as simple and sent it to Haiku. And here is the solution. Next, a slightly more complex task, but still not at the Opus level. Write a Telegram post plan about launching a new service. Five points. Sending it. And Haiku determines the complexity as normal and hands the task to Sonnet. And it is currently working on the answer. And this is the plan Sonnet prepared for the post. Now, let's give it a truly difficult task.
Design a database for a delivery service handling 50,000 orders. Create the tables, relationships, and indexes. Estimate the read and write load. And they send this task. Jeff decides it’s a complex task and hands it over to Opus. That’s exactly right. Only he can handle this. From this fairly simple example, you can see that Jeff actually handles distributing tasks to different models very well. Meaning you can build a really cool routing system with him.
And Jeff himself is, well, very cheap, so you’ll barely pay anything for him at all. But your monthly savings could be huge. Now let’s move on to the third, most interesting case. The previous two cases were simple enough to help you understand how Jeff works. And now, let’s move to a more complex one, where my agent will call Jeff instead of me. I built a skill for Claude Code called Jeff Decider. If the agent is sent a list of more than twenty entries , it won't look through them itself.
It will send this list to Jeff with questions and, after they pass the filter, receive a response back. To make it clear, let me show you with those same 300 requests. I wrote a prompt like this. In the file 'data', there are 300 incoming requests for our business. Find all urgent complaints where the client is unhappy and cannot wait, and write a draft response for each from a manager's perspective, three to four sentences long, in the file 'answer.m'.
Do not ask clarifying questions, and I send this prompt. And just so you understand, our agent will manually go through all these 300 requests, and it will take, well, quite a lot of time. While it reads all these 300 requests, let’s do the same thing with the skill. I’m sending the exact same prompt. And let’s see what happens. Just so you know, without the skill, the agent is still reading. And look, with the skill, it has already written all 25 responses, while without the skill, it’s still reading.
In the end, the agent without the skill spent 5 minutes reading all 300 requests and selected 37 urgent complaints. Then it wrote responses to them, but it took a really long time. And there are two buts here. First, it selected 37 complaints. Although, in reality, there were about 25 urgent complaints. So, it also selected them incorrectly. With the skill, however, it selected all 25 urgent complaints, wrote responses for them, and did it, well, much faster, spending only 1 cent on it.
So, using Jeff via API is very cheap. Plus, it speeds up the work significantly. By the way, you can also find the link to the Skill itself in my Telegram bot via the first link in the description. Jeff is a model that won't replace Fable or Astra for us, but it can work alongside them, which speeds up the work significantly. So, if you want to see more videos where I explore Jeff, be sure to write about it in the comments and subscribe to the channel so you don't miss out.
Also, watch my video where I practically live-code a working project for a business from scratch. Bye. Yeah.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Use this transcript
Three free tools that work on the material around a video like this one. No signup, no login.
Hook Analyzer
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Policy Pre-Flight
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Channel Skill Generator
Read this channel's public videos and transcripts, and download a writing brief for it.