Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Sharbel A. · @sharbelxyz
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in Sharbel A.'s most watched videos.
Most replayed moment at 13:32
3.9x that video's typical replay level
built-in memory, which is great. Here is how I would think about memory providers. Memo is interesting if you want a dedicated memory layer for personalized AI agents. It focuses on extracting, storing, linking, and retrieving memories efficiently. Their
Said at 13:25
Most replayed moment at 6:37
5.8x that video's typical replay level
trading strategy. Tests it. If it's better, it keeps it. If it's worse, it discards it and tries again. So, that's exactly what I built. Okay, here's the system we have at play. I gave it two years of crypto data,
Said at 6:31
Most replayed moment at 1:53
3.7x that video's typical replay level
let's install it together. Okay, so for step one, we need to actually start by installing Bullpen's CLI so that we can actually do everything that we want to do. And for that, let's first open Claude and put in the dangerously skip
Said at 1:47
The graph counts replays. It does not show where viewers stopped watching.
Words
9,159
Runtime
1:02:35
Speaking pace
146wpm
Reading time
38min
146 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Most people are testing AI models wrong. They ask is chat GPT better than Claude or is Gemini better than GPT or what is the best AI model right now? But that is the wrong question in my opinion. That is like asking is an engineer better than a designer better at what exactly? Because the model I use to build code bases is not the same model I use to write scripts.
73 words, the words spoken in the first 30 seconds at 146 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 783 |
| Average words per sentence | 11.7 |
| Longest sentence | 91 words |
| Questions asked | 91 |
| Sentences containing a number | 83 |
Most used terms
Filler phrases
149 in total: like 60 · uh 36 · I mean 22 · actually 11 · you know 7 · literally 5 · sort of 3 · kind of 2 · um 2 · right? 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
Most people are testing AI models wrong. They ask is chat GPT better than Claude or is Gemini better than GPT or what is the best AI model right now? But that is the wrong question in my opinion. That is like asking is an engineer better than a designer better at what exactly? Because the model I use to build code bases is not the same model I use to write scripts. The model I use for deep research is not always the model I use to brainstorm.
The model I use inside agents is not always the model I use as an everyday assistant. So in this video, I'm going to do this differently. I'm going to attach a task to each AI model. Who will be the best coding model, the best writing model, the best deep research model, the best reasoning model, the best creative brainstorming model, the best everyday assistant, the best cheap model, the model I was wrong about, the model most people should probably stop using, and the model I would pick if I could only use one model.
This is not going to be a surface level tier list. I'll show you examples from the kind of work I actually do, whether it's building code bases, writing content, researching markets, running AI agents, and operating a a day-to-day business. The real goal of this video is simple. By the end of this video, you should know which model to use for which job. Because the best AI users are not loyal to one model. They route the task to the right model.
Let's get into it. Here's the problem. Every week there's a new model, a new benchmark, a new leaderboard, a new ex post or thread or article telling you that everything changed overnight. And if you try to keep up with it like a sports fan, you will guaranteed to be losing. You become a claude person or a chatt person or a Gemini person and then every conversation turns into some sort of model tribalism. But in real work that is not how I use AI.
I run a marketing agency. We've done around $4 million in revenue in the last two years. And I use AI for real operational work. I use clot code for building and codebased work. I use Hermes for automations, workflows, scheduled tasks and operations. I use openclaw for agent experiments. And I use notion for planning and content systems. And across that stack of tools, the question is never just which model is smartest.
The question is what job am I hiring this model to do? I mean, if I need a code base changed, I need a model that can understand files, make safe edits, and run tests. If I need a script written, I need taste, and I need a model with a good voice. If I need research, then I need a model that's great at fetching and finding and compiling resources and synthesizing. If I need an agent, I need a tool discipline and verification.
If I need visual ideas, then I need range, creative direction, and the best image generation qualities. Those are, as you can see, completely different jobs. It's as though I'm going to my dentist because he's the best dentist in the world and telling him, >> "Hey, you're the best dentist in the world. Can you give me diet tips?" >> You can see how being the best at something doesn't really mean being the best at everything.
What I'm going to do in this video is walk you through the categories, the contenders, my pick, and exactly how I would test each one of the models. Starting with model number one. All right, let's start with coding because this is where the difference is the easiest to see. I assume my honest answer for best coding model right now is that it is very neck to neck between clawed opus 4.7 and clot code and codeex gpt 5.5.
Clot code with opus is still the coding workflow I find myself using the most throughout the previous month. But to be honest, this week I started seriously using codecs with GPT 5.5 a whole lot more and I'm starting to understand why so many builders are moving towards it. I mean it is fast. It feels strong inside the coding workflow and it is close enough that I do not want to make a fake confident claim before I have more time spent with it.
So instead of pretending there is one obvious winner, I will show the data then show the head-to-head test. Now I want to be specific about what I mean by coding. I do not mean asking a model to write a function in a blank chat window. This is not the real test. The real test is whether it can work inside a codebase without making a mess. Can it understand the repo? Can it make changes across multiple files? Can it run tests?
Can it debug from real error messages? Can it make the smallest safe change instead of rewriting half the app or the entire app even worse? That is why I still trust Claude most from my own history. But the data makes the codeex with GPT 5.5 side impossible to ignore right now. On LM arena webdev, claude opus 4.7 is sitting at the top while GPT 5.5 in a codec style harness is also in the top group on a gentic coding on the swench style comparisons.
Claude Opus 4.7 has a very strong public signal on code editing benchmarks like Ader open AI models have also been extremely competitive. So the real story is not claude destroys everything. The real story is claude code and codecs are now close enough that your workflow and test case matter a lot. So, if we can't decide using data, it's time to decide using this test. Now, here's what I'm going to do. We have clawed code on the left and Codeex on the right, side by side, and we're going to be running the same exact prompt with both of them, and seeing which one performs better, which one we will crown as the best coding AI model.
Now, here's what I care about before running this. I do not care which model writes the prettiest code in a blank chat window. I mean, I don't understand code either way. It's not like, you know, I I'd be a good judge at that if that was the test. I care which model behaves better when the task is real. I mean, this is exactly why I'm going to give it a real repo of mine. Does it understand the repo before even touching it?
Does it choose the right files? Does it make the smallest safe change? Does it run the tests before committing? And when something breaks, does it actually read the error or does it start guessing? That is the difference between a model that can code and a model I would trust inside a real workflow. So let's run this test right now by going with both. So trust workspace. Let's send it. And let's also give GPT 5.5 the highest um so that we're being fair.
Let's put them both on extra high. And here's the prompt I'm giving them. By the way, inspect my Xpilot GitHub repo and add a small new feature that improves the user experience. So I want to also see does it understand the goal here. I want users to get hooked on our app. So, I'm choosing to stay vague. I also want to test out creativity in this because people don't always have the easiest time verbalizing what is it that they want.
They know the goal that they want to get to. They don't necessarily know how they want to reach that. The rules that I've given it are 10 rules. Don't edit immediately. First, explain the architecture in plain English. Identify the files involved. Explain the safest implementation plan. Make the minimum safe change. Run tests. Type check lint or the closest available verification command. Show the difference. Explain what changed what you verified and what risks remain.
If anything fails, read the error and fix the specific issue. Don't just guess and do not use m dashes. Uh that's just a personal. So cloud code needs permission which I've just given it. Uh, another permission. I should have just opened the terminal or else we're going to be here all day giving it permission. Meanwhile, GPT has started. I'll start by locating the repo and reading only structure configuration first. Okay.
So, here is GPT. No edits yet. Here's the architecture in plain English. This is the Charbella XYZ dash or slashxpilot repo. So, GPT has chosen the files that it wants to edit. It has come up with the safest implementation plan. It wants to add a small personalized daily focus card to the dashboard home. And by the way, I think let me actually do this. Open a new window. And by the way, I do think it would be worthwhile for me to show you the app that I'm asking it to improve.
I'm asking both of those models to improve. This is it. It's called Founderfunnel. It's an app that we've developed for founders to be able to stay on top of all of the X- trends in their niche, happening in their niche, and helping them create more content that fits the current trend so that they're on topic, on trend, they can garner as much attention as possible and be timely. This is what it looks like. It suggests different content ideas.
It can help them write new drafts, discover what's trending, write an article, uh, so on and so forth. That's the gist of it. This is the trending tab of things that are trending in the last 24 hours. You get the gist. Let's go back to our two contenders. Okay, so both are moving. I can see. We'll just let them go on and I'll come back to check on them once they're both done or almost done. All right. And GPT codeex is done with their change.
And this is what they've shipped. This is the website. So, as you can see, empty here at the top. Meanwhile, this is the change that Codeex just shipped. It added this at the top of the website which it's calling a daily focus which funny enough Claude also wants to add a daily focus. So both of them for some reason chose to add this new feature called daily focus. uh and looking at it, I mean it is nice to see ship one draft today.
It's giving you something to focus on, but then again, that is the intent of the website is to help you create drafts. Maybe it's a smart idea to remind the user logging in that they need to ship drafts and that is the point of the entire website. I don't know. I'm I'm not too happy about this, but then again, Claude also is working on the same thing. So, I'm starting off a little skeptical. So, none of these buttons uh do anything.
So, this is our streak, which now there's two places where our streak shows up. We have our best streak previously, which we didn't have before, which is nice to see. Fastest win, which is not clickable. We cannot click it to see what could be our fastest win. And we can click review drafts. And it works. It opens our draft section. All right. So, this is GPT's uh version with two errors currently. Let's go and see Claude's version.
Okay. Okay. And I just realized it was a bad idea not mentioning to Claude that GPT is working on the same task at the same time because it literally just shipped the same exact thing. And when I asked it, I was like, "Wait a minute. I just saw this right now." He said that yeah, by the time he started, he saw that there was a file already in progress. So he decided to he thought it was me working on the code base and thought that was what I wanted.
So I told him to scratch it and start again. Assume that he's never seen this change happen. So we're going to come back once it's done and see what he ships instead. Whether he chooses to ship Daily Focus, another version of Daily Focus, or something else entirely. Let's see. Okay, so I told it this. It gave me a bunch of ideas. I told it, I don't want to give my input. Choose your favorite and let's test it out and it seems to have chosen to rework the daily focus function.
So, let's see its own version. Now, update daily focus. Get daily focus signature and branches and the card render. Let's see its own version. All right. And it's done. And to be honest, I really like this a lot better because what it did was take that daily focus idea and attach a draft to it so that it's choosing what it thinks is the top draft for this person to post. So, this person has 88 pending drafts, which can be very numbing as a founder, a creator.
Like, holy which which draft do I go with? Oh, which which of my ideas do I actually post? And it's suggesting what it thinks is the strongest draft for them to publish. And it kept the streak. It kept the best. It changed the button that did didn't do anything to how many drafts they have pending. So, it's a status. It's not a button either. But at least I personally think this makes a whole lot more sense for it to be at the top of the website.
A user logs in and the first thing they see is just one focus draft, just one daily focus. I really like this uh a lot. Okay, having run both tests, my honest answer right now is this. Claude Opus 4.7 and Claude Code is still the coding setup I trust most from repeated use, but codeex with GPT 5.5 is now close enough that I no one can ignore it anymore. I mean, the test that I've given them is very vague, but it's also vague on purpose.
I don't think most people when they go in to enhance their website, to work on their website, have an idea or verbalize well enough what it is that they want out of this, say, new update in our test example. This is why I kept mine very vague. I didn't request a specific request. I asked them to enhance the user experience and they both landed on daily focus. Claude opus was a bit more advanced, a bit more thought out if I can [snorts] put it that way.
But codeex has a lot of great features as well. I'm sure if I use something like slashgoal for example with GPT 5.5 with it being user experience or user retention goes up it would have gave me something with a lot more oomph with a lot more fire but I wanted to give them a leveling a leveled playing field and that is where we're at. I have been testing GPT 5.5 a lot this week and honestly I understand why a lot of people are moving towards it.
I mean it is fast, it is sharp and for some task it feels absolutely neck and neck with clot code. So and I would not tell you codeex is the new king and it has completely replaced clot code after just one week of me testing it. The answer is a little more useful than that. If you're doing serious repo work, test clot code and codeex side by side on the same task and give them the same repo. Give them the same feature.
Force both to plan, edit, test and explain the difference. Then judge the workflow for yourself, not just the final answer. For me, that is the coding category right now. Two models at the top. Clot Code with Opus 4.7 is my trusted pick so far and Codeex with GPT 5.5 is the challenger that might earn the slot very soon. But coding is only one kind of intelligence. The next category [snorts] is where the technically smartest model can still lose because writing is about taste.
For writing and content, my personal pick is Claude, specifically Claude Opus when I want the highest quality and Claude Sonnet when I want fast day-to-day writing and editing. The runner up is also GPT 5.5, especially when the writing also needs strong reasoning or strategy. But for voice scripts and content, Claude is still the model that I use the most when it comes to that and that I trust the most. Here's why. Writing is not just producing words.
Writing is taste. It's knowing what to remove. It's knowing when something sounds fake. It's knowing when a sentence is technically correct but emotionally dead. A bad AI model will write something like this. Artificial intelligence is revolutionizing the way we work. And in today's video, we're going to explore the best AI models for productivity. [laughter] That's this just sounds wrong. It's just useless. It sounds like every AI video ever made.
What I want is something sharper, something that sounds like me or something that I can riff over. something like most people are using AI models wrong. They pick one model, use it for everything, and then wonder why half their outputs feel wrong or average. And that sounds a whole lot closer. I mean, to me, that is how I want my agent to sound like. I want him to understand tension. I want him to help me voice out a point of view.
I want it to sound like a person. But without further ado, let's give them a test. Okay. And for this test, I'm going to be testing out GPT 5.5, Opus 4.7, Sonnet 4.6, Gemini Pro uh 3.1, and Kimmy K2 as well. Kimmy K2.6 to be specific. And I'm going to give them the same exact prompt. Rewrite this YouTube intro and my voice. The goal is for the viewer. The viewer clicked because they want to know which AI model is best for each task.
The intro must confirm that fast. Then make the deeper point that the real skill is model routing. My voice. I've given it instructions. Bad intro to rewrite. So I've given it the intro I wanted to rewrite. And I've given it some rules as well. And let's start with GPT 5.5. The first option that it gave me is to start with the mistake first. You're probably using the wrong AI model for half your work. Not because the model is bad, because you're asking one model to do everything.
ChachiPT, Claude, Gemini, Grock, they're not interchangeable. Some are better for coding, some are better for research, yada yada. So, in this video, I'm going to show you which model I'd actually use for each task. I really like this uh intro, this hook. right here. I'm a fan of it because he's he knew how to give stakes. He knew how to grab attention, start with an objection, a contrarian take, and then he knew how to come through with a resolution in the same intro, which is very hard, very difficult to do for most copywriters, let alone an AI agent.
Uh, option number two, it gave me operator workflow angle and option number three, caveat and tension angle. There's no single best AI model. That answer is annoying, but it's true. Not too much of a fan of option three. Option two, most people compare AI models like they're picking a favorite app. Sounds very AI generated. Option number two. Option number one is is great. I really like it. So, GPT 5.5 starting off really strong.
And the way I'll be ranking them, by the way, is on six different metrics. Number one, does it sound human? Number two, does it sound like me? Does it actually sound like my voice? Number three, does it start with tension? Does it that hook have tension? Number four, does it use filler words? A lot of filler words. Number five, does it sound like a just AI guru, AI slop in any way? And number six, would I use it? Could I see myself using it?
So, let's start with GPT 5.5. We're going to rank each on a scale of one through five, then sum up the scores and see which one wins. For GPT 5.5, I'm going to give it for its option one, which is the one that I'm picking as its best one. I'm going to give it a strong four. Does it sound like me? I'm also going to give it a strong four. Does it have tension? I think so. I'm going to give it a five out of five on that.
Does it use a lot of filler words? Not really. So, I'm going to give it a strong four here as well. Does it sound AI? I mean, I would probably tweak a couple of things from that hook, but it is in my opinion 90% of the way there. So, I'm going to give it another four here. Would I use it? Absolutely. I'm going to give it a five. So, in total, GPT 5.5 is sitting at a 26. 26 out of 30, which is a pretty damn good score.
All right, let's move on to Opus 4.7 with its three options. Option number one, most people pick one AI model and try to force every task through it. That's the mistake. Ah, sounds very AI AI generated. Well, I mean, it is, but unfortunate. Option number two, if you're paying for one AI subscription and using it for everything, you're doing it wrong. I'll get to the rankings. Best model for writing, best for code, best for research, blah blah blah.
Also, not a big fan of it. Uh, option number three, there's no best AI model. There's a best model for the thing you're doing right now. I really like this. There's no best AI model. There's a best model for the thing you're doing right now. I think I'm going to go with option number three as my personal favorite from it. Again, I mean, this entire ranking is subjective, but you can go ahead and rank them out yourself as well as we go through them.
So, for Opus 4.7, does it sound human? I'm going to give it a strong four here. Does it sound like me? I'm also going to give it a four. Uh, does it start with immediate tension? Absolutely. I really like the way they it started this hook. Does it use a lot of filler sentences? Honestly, no. I might give it a five here because it's straight to the point. It's a lot shorter and more concise than GPT's draft. So, a lot less need to edit it, if that makes sense.
Does it sound like it's an AI guru? I mean, there are some giveaways. Not picking a winner. I'll show you the workflow I run. I'm going to give it a four here. A strong four. And would I use it? Absolutely. That's a five here. Which brings the total if I put the numbers up to 27. So just one point more than GPT 5.5. So the two aren't that far apart when it comes to writing. Let's look at Sonnet 4.6. A lot shorter, which I do like.
But let's see if there are any if they are any good. Option one is the mistake angle. Most people pick one AI model and use it for everything. That's the problem. In this video, I'll tell you exactly which model to use for which task and more importantly how to stop making that switch manually. Oh, that is that is a very nice hook. Option number two. Here's the short answer. Clawed for writing. Hold on. What is that?
A intro? That feels just like somewhere in the middle of a video. Option number three, the operator angle. You click this because you want a tier list. Fine, I'll give you one. But the people actually saving time with AI aren't paying me. Okay, I like this, but it's not my favorite option. One, I'll go ahead and lock it in as my favorite and grade it. Does it sound human? Uh, where are we putting this? That's the problem.
Does it sound human? I'm going to give it a four. Not a strong four, though. I'm not going to give it like a 3.6 or something. I'm just going to put it at four. Does it sound like me? Uh, I mean, it could sound like anyone. It doesn't really have personality, I feel. So, I'm going to give it a three on this. Does it have tension? I've given the other two models a five on this. I think Sonnet gets uh four on this. Does it use a lot of filler words?
None. It just goes straight to the point. So, a five on this. Does it sound AI generated? I'm going to give it a weak four, which is still a four. And would I use it? Honestly, I could see myself using it, but it does need some editing. So, I'm going to give it a four on this, which brings the total to it brings the total to 24. So, so far we have Opus 4.7 at 27 points. GPT 5.5 as a runner up at 26 points and Sonnet 4.6 at 24 points in third place.
Let's go ahead and meet our next contestant, which is Gim Jim. Gimini. Gemini 3.1. Gimini. That is a nice new name for Gemini. Gimini. All right, let's see what Jiminy has in store. Option number one, mistake first. Most people are picking AI models backwards. They ask, "What's the best model?" Wrong question. The real question is, uh, sounds very sloppy. Uh, I have to say not a fan of this one. Option two, workflow angle.
If you're using one AI model for everything, you're doing it wrong. Not because the model is bad, because the task changes. It doesn't give me uh really that that strong oomph that I'm looking for. Third one, there's no single best AI model. That's the part most comparisons get wrong. I think this is its best hook out of the three. Still not a banger, but let's go with that one. So, on a human scale, does it sound human?
I mean, I'm going to give it a week four, same as I did for Sonnet. Uh, does it sound like me? Not really. It doesn't sound much like anyone. Does it start with tension? I'm looking through all of its generations and I don't really feel tension. like that tension that just grabs you, makes you want to watch more. So, I'll give it a three there. Does it use a lot of filler sentences? I mean, yeah, it it pretty much sprinkles fillers throughout.
So, I'll give it a three here as well. Does it sound like AI like an AI guru? I'm going to give it a three here. Which by the way for for the filler ranking and for the AI guru ranking a high score doesn't mean that they sound AI generated. It means that they sound the least AI generated. So it's sort of a backwards ranking. And would I use it? I'm also going to give it a three. I don't see myself using this copy. Which brings us brings the total to to 19.
All right. So quite far from the rest of the competition at 19. Let's have a look at our last option. Kimmy K2.6. The first copy is most people are using AI models wrong. They ask which one is the best as if there is one answer. There isn't. That's not a bad copy. A lot of room for edit, but not a bad copy. Chat GPT might be the best place to start a messy idea. Claude might be better when you need long context and clean writing.
This one sort of gave you everything you need to know about the video before the video even started, which is a little too much detail. A lot of detail. Let's look at the other two. If you're trying to find the best AI model, you're already making the wrong move. That sounds harsh, but it's true. Here is the simple answer. Different AI models are good at different jobs. Okay, this sounds like it's somewhere in the middle of a video.
Not really a hook. I think option one is its strongest hook. So, let's grade that one. Does it sound human? I mean, it does. I'm going to give it somewhat of a four. Does it sound like me? I mean, sure. There's a lot more skin to this hook. I just don't like the overall structure. So, does it have tension? I'm going to give it a three here because I feel like it kills the tension. It doesn't really leave a trail behind.
In in copyrightiting, you always want to leave uh suspense, room for suspense. Oh, what's going to come next? What's going to come next? With Kimmy's intro, it just gives the entire video away in the first 5 seconds. Filler. Does it use a lot of filler sentences? I'm going to give it a three here. Does it sound AI generated or an AI guru? Is it all wise and knowing? I'm going to give it a four. I'm not going to be too harsh on it.
Would I use it? I'm going to give it a three. Not really. Just very mid. So, let's come up with the total. And the total for Kimmy K2.6 is 21 points. That gives us our winner for this category with Claude Opus 4.7 at 27 points and very close to it the runner up GPT 5.5 at 26 points. And with that we have our winner for the writing section and it is Claude Opus 4.7. Two wins in a row but can it keep up? Honestly, I'm very surprised with GPT 5.5 because I'm finding that it can be extremely strong for things like strategic thinking, but Claude tends to be better at preserving voice and removing that AI smell.
I feel like where it wins is with scripts, positioning, editing, voice matching, content angles, and narrative structure. It really has a knowhow of good copywriting and so does GPT 5.5. That was a really good copy. I don't think there are any close contenders to those two when it comes to writing specifically. Whether it's to help you with YouTube, Twitter, uh brainstorm content, come up with ideas, ideiation, those are very two strong picks.
I don't think honestly you can go wrong with either of these. With that, let's move to category number three. For section three, we're going to be covering deep research. And my personal pick for that, the model that I find myself using the most is Gemini 3.1 Pro or Gemini 3 Pro. The runner up here is GPT 5.5 which I also tend to use a lot for that. And as much as I give a hard time for Gemini, uh it is one of the places where I think a lot of people underestimate Gemini.
If the task is read a lot of information and find the signal, Gemini deserves a real spot in the workflow. I'm talking about long documents, PDFs, market research, competitor research, tool docs, things like big context windows, sourceheavy synthesis, synthesis. This matters because research is where AI can be the most dangerous. A model can sound confident and still be completely wrong. So for research, I don't just care about the final answer.
I care about how it separates evidence from opinion. And when it comes to Gemini, it's a really affordable model for what it can do and especially compared to its competition. But that being said, let's test it out. Is there a better model out there? Let's see. And for this test, I'm going to be giving all different agents this prompt. Read the attached PDF and help me improve a YouTube script. The context the YouTube script the video sorry is called I tested the top AI models blah blah blah.
Here's the promise. The return is or rather return for me the core framework of the PDF in plain English. The strongest ideas what applies directly to this AI model comparison video. What would make the video feel generic? the 10 specific edits I should make to improve retention, a better opening structure for the first 60 seconds, and a checklist I can use before filming. I've given the different models the same exact script, but I've included a pivot, a caveat.
I've asked them here at the bottom to score themselves on a scale of one to five on each of these source quality, citation accuracy, synthesis, currentness, contradiction handling, and actionability. I'm going to be going through all of their answers, and also through all of their scoring. I'm especially interested in seeing what they rate their currentness because I've given them a file. In my prompt, I didn't even ask them to check sources.
So, this should come back as very low on their own self-coring here. What I'm really testing is honesty. Are they hallucinating? Are they being honest about the sources they're using? and this is the best way to test it out. So, without further ado, let me go ahead and read a lot of what they've given me and give them each their rankings, starting with Opus 4.7. All right, so I just went through Opus' results and honestly, it really surprised me in a very good way because let me show you.
I mean, throughout this entire thing, it includes citations. It gives us literally a structure that we can follow through. It gives us 10 clear points that are actionable that it came out with its research. And when it came to honesty, it shocked me how it created itself. It said source quality. I pulled this research from one source. Citation accuracy, no citations because no sources. Every claim I made was from that PDF.
And currentness, one out of five. I have no idea what the current model landscape looks like as of today. I didn't research. I only used the PDF. It was very honest and it came back with a very honest scoring system as well. For that, I've given its quality a five. citations uh a five or rather now that I think about it or I say it out loud more so of a four synthesis a four current which really the ranking here is honesty I've given it a five it was super honest contradiction a four actionability a five which brings its total score to 27 let's look at the next contender for GPT 5.5 Five.
I've noticed that, and this is something that I have against GPT as a whole from the beginning, is it's always so confident. It could be giving you BS and saying it so confidently and so sure of itself. I ranked it. The ranking has come out high. I'll save you the hassle of going through all the scores. But one thing I want to note is its currentness. It first gave me a currentness of 3.5 and I was like, wait, are you talking about the above?
Are you sure 3.5 out of five? And then it changed its answer to two out of five, which isn't true. It should be a one because it's not current. It's only verifying through that PD, that one PDF file. It has done zero uh online research with this. What I wanted is it's its ability to extract information from that document. And for that its total score comes out to 25. 25 out of 30. So Gemini I just went through its answers and surprisingly it was very honest.
It gave its own currentness almost none almost non-existent. It did not check what the current top models are so on and so forth. It came out to a total score of 25, which ties it to GPT 5.5 with Opus being in the lead. Let's see. Can Kimmy K2.6 win this one. For Kimmy K2.6, I have two things specifically to note here. Number one, it did not really do well at citing the source. Uh, that is something I noted down. and a place where it lost some points.
The second one is contradiction is it contradicted itself and I find myself also seeing that in my day-to-day uses. It tends to contradict itself with its sources. Uh it will fetch a source from here and a source from here and you know mash them both together but each one is giving you a completely different uh point of view. And lastly, something interesting. It's honesty, which we ranked as currentness. It said not available.
So, it decided not to rank itself or a three out of five. Like, come on. Which one is it? You can't have both. It was somewhat honest, but it didn't change its score to reflect that it did not it it is not current. For that, I've given it a four out of five on that metric. but its score came down as the lowest at 22. With that, we have our winner, Opus 4.7. Yet again, can it still remain undefeated in the next round?
Let's look. If I had to pick the best everyday assistant, I'd pick GPT 5.5. The runner up for me would be Claude Sonnet 4.6. And that's what I find myself using every day, that alternation. those two together. An everyday assistant is different from the best model overall. This is the model you open when you need help quickly. Explain this. Plan this for me. Help me rewrite this. Compare these options. Summarize this for me.
Help me think. Draft this message out very quickly. For that, I want a model that is broadly strong, flexible, and good at moving between tasks. GPT 5.5 is my pick because it is strong enough across reasoning, planning, coding, explanations, uh, business questions, and just general work overall. Claude Sonnet is the runner up because it is fast and it's practical and it's very strong for a lot of daily work especially writing and agent style tasks.
But with that, should I be changing any of them? Let's test them out. For this, I want to test them out by giving them a task that is very realistic and something that I would actually ask my AI agent to give me on a day-to-day basis, which is I have 45 minutes before a meeting. Here are the notes. I gave it a bunch of meeting notes. Turn these notes into a short meeting agenda. the three decisions we need to make, the risks I should bring up, the strongest recommendation, and a follow-up message I can send after the meeting.
I'm going to be ranking the models on accuracy, speed, helpfulness, personality, those four ones. Speed is very important for your everyday assistant. Accuracy is extremely important. You don't want it hallucinating. Helpfulness. How much has it been helpful? How much do you find it helpful? How much does it helps help you point things out that you didn't notice? And personality, you want to be speaking to someone that feels like a person, not to the most AI agentic creature in the world.
So, with that, I'll go ahead look through all of their different responses and come back with a ranking. All right. So, this turned out to be somewhat expected, but also in a lot of ways interesting. And I mean what surprised me first was the speed at which all of the different models replied to me. I feel like all of those models, especially throughout the last year, have really had massive improvements when it comes to speed.
Literally, I think all of them responded to me at the same time uh on Telegram, but also the ones I used outside of Telegram, they were really quick. Second was accuracy. I mean, all of them summed up the notes I gave them in a very good way, in a very accurate way, although the helpfulness is where a lot of them varied. with GPT. He gave me recommendations and things to focus on that were very concise and very straight to the point.
What is the next video? Pick one idea, not four half ideas. What demand signal is strong enough to justify filming? With Kimmy and Gemini, their recommendations were a lot longer and a lot more, I don't know, it felt corporate. So they lost points on helpfulness and personality as well. With that the rankings are as follows. Gemini with 15 points. Kimmy with 14 points. So just behind Gemini, GPT with 19 points overall.
And Sonnet 4.6 6 with 19 points as well. So that's literally a neck between the two. What I loved about Sonnet, by the way, is how concise it came back. Now granted, I am testing GPT 5.5 via Hermes, which has its own personality and you know some preferences that I've customized. So I have not redacted points from it for that for uh conciseness in the overall text. I did redact points however for conciseness in the uh specific sections.
So with that we have a tie. Our first head-to-head tie between GPT 5.5 and sonnet 4.6. All right. one of my favorite categories because this one is not just about what is the best model. It's what is the best model for its value. What is the best cheap model? And the best cheap model is not the model that feels impressive in a demo. It's the model that does simple work fast and cheaply without needing the smartest brain in the room.
For this category, we're going to be pinning up and testing Gemini Flash style models as well as models like Quen and Kimmy. And the reason for this category is simple. Not every task deserves your best most expensive model. If I'm doing high volume summarization, classification, uh, metadata cleanup or common grouping, I don't want to pay for the most expensive reasoning model. That is just wasteful. So, without further ado, let's test it out.
This is a very interesting round because I just went through all of them and I realized that each of the cheaper models have a lot of edges and a lot of faults that are very different from one another. Let me give you an example because Kimmy in Kimmy's case he was the slowest model out of the ones I tested. Now, when I say slowest, I don't mean like it took 10 minutes instead of 1 minute. Each of them took under a minute to respond back, which is fast.
But one of the first ways I ranked them was on speed, and Kimmy was the slowest one out of the three. The second way I scored them was accuracy where Gemini just summarized a lot more and was not too accurate with his summarization. He just bundled a lot of the comment sections up. And the third category was helpfulness, which all of them scored about equally the same. And all three of them came out with a score of 14.
I don't know if you can hear it, but some guy is trying to sabotage this ranking. He was sent from either OpenAI or Anthropic to sabotage this entire ranking. But surprisingly, the three models, those three models tied at 14 all together and neck to neck to neck. So what I did was since this is the cheapest best model category, I decided to compare cost and make that tiebreaker be the cost comparison. And this is where Quen came out as the cheapest.
Hence why it is winning the first place, Gemini as the second cheapest. Uh which is where Gemini comes in at second place with 18 points. So Quen 19 points, Gemini 18 points, and Kimmy 17 points as the third cheaper. Again, all three of them are insanely cheap, and they each have different strength in different places, but I'm honestly very surprised at how good they all were when you look at everything that comes with them.
And this brings us to our last test, our last competition, which is our creative competition. For this one, I'm going to make the creative decision of limiting this test to models with image generation because that is really what I want to be testing here. This test will be in two parts. Number one, concept design. So, I'll be sending them two prompts. The first one will be around the concept. Can it come up with a good YouTube thumbnail concept?
And the second one is the actual image, the one that matters, which is can it come up with a really good image, a thumbnail that I would actually see myself using. So, let's see. We're going to be pinning the top three models in that category in my opinion, which are GPT 5.5 with their latest images 2.0 O release, Grock, which is great at images, but can it be as good as the rest? And finally, Gemini's Nano Banana, which is insanely good.
But the question is, which one is the best? Let's find out. All right, for this final test, I'm going to be ranking each model on four different scores. Number one is the concepts that they come up with. How good are they? Number two, the actual image. How good is it? Number three, does it feel too AI generated? Or in other words, does it change my features a lot because I'm going to be using my face and seeing how close to my face does it actually come out with the thumbnail with?
And lastly, number four is would I actually use it? Would I feel like, you know what, that's a thumbnail I'd use without editing. How close is it to that? So with that being said, Gemini, this is Gemini, right? Has strong scored very strongly on all of them to be honest. Uh except for the concept and it aligns with how I find myself using Gemini already. I find myself a lot of the time for my YouTube thumbnail creation going to GPT getting the concept from GPT and then going to Gemini and having it create the image for me which is very interesting and I've seen this with the concept that it came out with.
I asked it for 20 concepts and a lot of them were good. I've given it a four out of five. So it's not the weakest. Grock, on the other hand, I felt when it came to concepts was the weakest because I didn't feel like it understood how detailed a concept has to be and how well and clean a YouTube thumbnail concept has to be. It really just gave me, you know, bland, quick answers for concepts. It felt like it was just rushing to come up with an answer. like, you know, that kid like doing an exam that just wants to be out of the room as soon as possible.
Whereas GPT scored the highest. He really went intentional with each thumbnail and outlined a layout, his own ranking. None of the other models did that, by the way. which ones it thought it was the best or were the best thumbnail concept along with facial expressions how they should be what are the colors that should be there why would it get clicked that was extremely well thought out and put together for that GPT gets the highest score on the concept level then we move to the actual image generated so for GPT this was the image it generated which shows a a confused face which shows a lot of the different models along with the right logos.
By the way, this is something a lot of people who use AI image generation. Something you guys will understand is sometimes you tell it, you know, put GPT's logo and it will just put the wonkiest logo out there. So, I'm very happy with that. for that it has uh one of the highest scores out of the models. When we go to Gemini, this is what it gave me. So you can see it repeated a lot of the logos. Now granted with Gemini, I find myself having to give it more instruction, more specificity for me to reach what I want, but I can see its intentions with the way it chose to do the thumbnail.
Overall, it's a well-rounded thumbnail. The facial expression is there. Um, it does look like me. It didn't distort me too much. With Grock, I really like the ambiance that it went for. Very dark, mysterious. I don't know if that is a clickable thumbnail, though. It I feel like it lacks depth and color. I can see what it was going for. I don't know if this is a thumbnail people would click on and a thumbnail that would perform on YouTube.
So for that, Grock had the lowest score from on that category. And when it comes to would I use it, it also has the lowest score compared to both Gemini and GPT. But with that, the final rankings are in third place. We have Grock, which is amazing at creating images, but not to the level of his competitors yet. In second place, and this is going to be shocking, I'm going to pick Nano Banana in second place. So, Gemini's Nano Banana in second place with 18 points out of 20, which is very high.
It's an amazingly good model. It's the model I find myself using the most when creating images, but it's the model I feel tends to be reliable on another, which is GPT, to come up with the concepts and help me really define that thumbnail in the way that I want it, which that dependency is a bit of a crutch for it. Uh, and that's what I found through my own personal use as well. So the way I thought about this ranking and a number one being GPT with images 2.0, it's a model where you don't really need anything else.
You can just stay on it. You can use it from concept ideation, brainstorming to execution all the way through. Granted, I prefer some parts of Gemini's image generation, but if I was to only use only use one model for image generation and one model only, it would have to be GPT. Look, when it comes to the best model, there's a lot of things that you need to consider. I mean strategy, technical explanations, cost, efficiency, coding, helpfulness, assistfulness.
Is that even a word? Assistfulness. How assisting it is. But the thing is is that matters only if you ever end up using just one model. But the whole point of this video is that you should not only use one if you don't have to. The model routing system is this. And our winners for this challenge are as follows. To recap, for coding, we had Claude Opus 4.7 come in first place with GPT 5.5 being literally neck and neck with it.
So those two models crush when it comes to coding. When it comes to writing and content creation, Claude Opus 4.7 and GPT 5.5 also dominate here as well. When it came to deep research, Claude Opus 4.7 came on top. When it comes to your everyday assistant, GPT 5.5 came in first place with Claude Sonnet 4.6 coming in very, very close as a very close second. When it comes to the best cheap model, Quen 3.6 came out on top in terms of cost, efficiency, speed, and finally, when it came to visual work and creative work, GPT Image 2.0 came in at first place.
And it's interesting when you look at this list right here because we don't have one model dominating every category, but rather a routing system of multiple models for the best performance overall. The big shift is this. Stop thinking of AI as one assistant. Start thinking of it as a team. You have the engineer. You have the writer over there. You have the researchers. You have the strategists, the operators, the designers, the cheap intern for repetitive tasks. >> Why did you have to single me out? >> And your job is not to be loyal to just one of them.
Your job is to route work accordingly. That is how we use AI in my team at my business. Clot code we use it for serious building. Hermes for workflows, automations and operations. Open claw for agent experiments. GPT 5.5 for image creation, concept design, my everyday assistant. Different models for different jobs. And the best part about it is if a new model launches tomorrow, which it very likely will, I don't need to panic.
I just run the same tests and I try to see does it code better, can it write better, can it research better, can it reason better or can it operate inside the tools that we use better or can it create better visuals? If the answer comes out as yes for any of those, then it earns a slot in my operation stack. If not, then maybe it's just hype or maybe there's a better alternative out there. Anyway, I hope you enjoyed this video because I had a ton of fun filming it.
And if you're building with AI agents or codeex or clot code or Hermes or OpenClaw or even trying to use this stuff in your everyday business, make sure to subscribe to my channel because I have a lot more content around the same exact niche. Anyway, I'll see you in the next
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.