Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Search · @theAIsearch
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Search's most watched videos.
Most replayed moment at 10:28
3.1x that video's typical replay level
and then fix any errors that it sees. And so afterwards, it rendered the video successfully and that's pretty much it. In just one prompt, I didn't even need to prompt it further. Here's our final result. >> Four companies, one quarter, and a half-trillion-dollar bet on artificial
Said at 10:22
Most replayed moment at 15:21
4.1x that video's typical replay level
and deep agent for only $10 a month. This is way cheaper than if you paid for each tool separately. Definitely check out chat.llm that comes with deep agent in the description below. You can think of the residual connections not as a simple pipe carrying the signal forward,
Said at 15:15
Most replayed moment at 12:41
6.0x that video's typical replay level
other content, Higgs Field is a game changer that will supercharge your production workflow. Try it today using the link in the description below. Now, if we dive deeper, here's how it works in technical terms. They used something called a Markov head. In probability theory, a Markov process assumes that
Said at 12:33
The graph counts replays. It does not show where viewers stopped watching.
Words
6,192
Runtime
35:31
Speaking pace
174wpm
Reading time
26min
174 words per minute, between the 160 25th percentile and the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
The Orange Sphincter is back. Claude Fable 5.1 is here, and this is currently the best AI model in the world. At least that's what they claim. In this video, I'm going to put it to the test so you can see for yourself what it can and cannot do. Plus, we're going to go over its benchmarks, specs, and limitations. If you're thinking of subscribing to Claude, definitely watch my video first because there are a ton of things you need to be aware of. Let's
87 words, the words spoken in the first 30 seconds at 174 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 489 |
| Average words per sentence | 12.7 |
| Longest sentence | 61 words |
| Questions asked | 2 |
| Sentences containing a number | 117 |
Most used terms
Filler phrases
89 in total: like 48 · actually 24 · basically 5 · I mean 4 · you know 4 · kind of 3 · right? 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
The Orange Sphincter is back. Claude Fable 5.1 is here, and this is currently the best AI model in the world. At least that's what they claim. In this video, I'm going to put it to the test so you can see for yourself what it can and cannot do. Plus, we're going to go over its benchmarks, specs, and limitations. If you're thinking of subscribing to Claude, definitely watch my video first because there are a ton of things you need to be aware of.
Let's jump right in with some tests. Now, as with all the Frontier models out there, they are designed for agentic tasks. That means you can give it a goal and it can autonomously use multiple tools and programs or even spin up sub agents and keep working on a task until it achieves your goal. And the best way to use such agentic models is not actually using this online chat interface, but instead using a harness. And the best harness to use for Claude is their own Claude code.
It looks like this. And this allows you to set up multiple projects and you can get agents to work on multiple things within each project at once. You can also give it access to local files so it can work with stuff straight from your computer. So for most of my demos in this video, I'm going to use Fable 5.1 in cloud code. All right, first let's test its spatial reasoning and its ability to generate 3D models and design different environments.
So, for my prompt, I'm going to write, "Create a fully furnished 3D interior design based on this floor plan image of an apartment." And I'm going to ask it to include full textures for everything. Give me three different designs. It should look extremely detailed, professional, and realistic. It should load efficiently on a web browser. I'm going to select Fable 5.1. And for the effort, I'm going to set this all the way to ultra code so you can get a sense of the best it can do.
Let's press run. All right. So here it decided to use 3JS to create the 3D model and then it worked for around 30 minutes. Now the original prototype was pretty good but it looked pretty basic. So I asked it to make everything even more detailed, realistic and impressive. And you know the crazy thing is after just like one prompt it already reached the freaking 5hour limit for Fable 5.1. It's probably because I set it to ultra code.
So after waiting like 4 hours, I told it to try again and then it continued and that was it. So after like two prompts, here is the final result. This actually looks really good. So first of all, if I show you an overhead view of the apartment, it does look exactly like the floor plan that I uploaded. All the furniture and other objects are exactly aligned with the reference floor plan. So this has pretty good spatial understanding.
Let me just rotate this scene around so you can see what the room looks like. So, here's what the bed looks like. And then here's the desk. Here's the grand piano. It was even able to create a pretty decent looking grand piano. And then let me rotate this so you can see the other side. So, here are the sofas and the lounge. And also, here is what the bathroom looks like. All right. So, that was Noir walnut. Next, the next design is Nordic Light.
So, it looks like this. Again, the arrangements of all the furniture and other objects are exactly like my reference floor plan, which is great. All right, so that's Nordic Light. And then next, we have Emerald Deco, which looks like this. So, you know, pretty good for two prompts. It's able to create a pretty detailed design of the apartment with decent looking objects and furniture, and it's actually really hard for even the Frontier models to understand floor plans like this.
So, this has very good spatial understanding. Now, for your reference, here is a detailed breakdown of the usage stats for this prompt. All right. Next, let's test its physics understanding and ray tracing capabilities. So, here's my prompt. Develop a ray tracing simulation featuring one sphere, one cube, and one pyramid floating in an infinite ocean with a blue sky. For each shape, add adjustable parameters like position, reflectivity, roughness, transparency, and other material properties.
For the ocean, add sliders for wind direction, wave size, and other settings. Also, add adjustable parameters for the sky. Put everything in a standalone HTML file. And here is the most important part. Do not use 3JS or any external libraries. It needs to code up everything from scratch. So, I'm really testing its inherent understanding of physics and lighting. And this is way harder than my previous ray tracing tests because here I'm also including the physics of an infinite ocean and different light settings from the sky.
Again, I'm going to set this to ultra code and then press run. All right, here is its response. And this worked for around 30 to 40 minutes, but that's pretty much it. It actually aced it in just one prompt. So here's the result. All right, as you can see, we have these three shapes floating in an infinite ocean with a sky. And it created all of this from scratch without 3JS or any external libraries. First, let's play around with the settings for the sphere.
So, the position of the sphere works. The size also works. Let's set the color to white. And then let's play around with reflectivity. So, let me adjust the slider. Here is what maximum reflectivity looks like. Here is zero reflectivity. And then let's also play around with roughness. So, roughness seems to work well. Let's drag this to zero. And then let's play around with transparency. So transparency seems to work pretty well as you can see here.
And then refraction IR that also seems to work. So here's what happens when I play around with the slider. All right. And then next we also have metallics. So here is full metallic. All right. Next let me play with the cube settings. So let's also move the position. Let's adjust the size. So the size works. Let me set the color to something like red. Reflectivity also works. Roughness. And these other settings also work.
And the pyramid also works. So there you go. All the settings from these three shapes work. Next, let's play around with the ocean. So here's what happens if we adjust the wind direction. Let's make the waves a bit larger. So that is what happens. Let's also adjust the wavelength, which seems to work. Next, let's adjust choppiness. It's a really subtle effect. Next, let's also adjust wave spread. So this is what happens if we decrease it.
Here's what happens when we increase it. And then here is wave speed. Very nice. Let me decrease it. We can also add foam here apparently. So let me increase the foam which looks like this. Next, let's play around with the sky and the sun. So we can also control the zenith color and the horizon color. So let's set this to something like red. Let me set this back to blue. Next, let me adjust the sun azmouth. And here is what happens.
Next, let me adjust the sun elevation so we can actually see the sun over here. We can also adjust the sun color. So, let's make this red. Let me decrease the sun elevation so it looks like it's sun setting. We can also increase the sun intensity. So, here is full intensity. Here's the lowest intensity. We also have the sun size. So, let me increase this and let me decrease this. And then here's the haze setting. So, here's low haze and high haze.
And then we have a ton of other settings here like cloud coverage, cloud scale, and also cloud speed. And then we also have several camera and render settings which also work. So I mean in just one prompt, it's able to render this pretty complicated ray tracing scene completely from scratch without any external libraries. So I am very impressed by this. The 3D rendering capabilities of Fable 5.1 are indeed state-of-the-art.
Next, let me pull up the usage stats for your reference. Now, the Frontier models out there are really good at using different interfaces. So, let's test out an example of its 3D modeling capabilities using Blender. Here's my prompt. Use Blender MCP at this address. Note that for my Blender, I I've already installed this Blender MCP. So, here I just need to click on connect to MCP server so that I can connect any AI agent to this local port.
And I'm going to get it to create an X-wing fighter spaceship with realistic texture and motion. make it incredibly detailed and faithful to the original spaceship. So, here's a time lapse of it actually pulling up Blender and creating the X-Wing fighter directly in my interface. Afterwards, it's also taking screenshots of the model and then verifying that everything looks correct. And here is our finished result. Let me just pause this first so we can inspect the spaceship.
You can see the details still aren't really realistic. It did add a cute little R2-D2 at the top here. But I mean, the overall design of this is quite basic. Here is the wireframe view. So, you can see it has generated all these different parts. And then here's the solid view. Here's the material preview. And then here is the final render. Honestly, I'm not too impressed by this. I actually prefer the generation from Opus 5 a bit more, which looks like this.
So, in terms of creating 3D assets or designing video games, front end, etc., It feels like Opus 5 is actually better than Fable 5.1. Next, let me pull up the usage stats for your reference. Now, the Frontier models are incredibly strong at agentic coding and tool use and autonomously carrying out a task using various platforms and interfaces and tools until it achieves your goal. So, here's a test of its capabilities.
First, I'm going to get it to search the web for the most recent earnings reports of Alphabet, Nvidia, Amazon, Apple, and Meta. And then let's get it to create a professional motion graphics presentation video that thoroughly compares the financials and future outlook. I'm not even going to tell it what tool it should use to create the video. It can decide for itself. Now for the voice over, because Claude can't generate a voice over itself.
I'm going to ask it to use Gemini TTS. Now the video should be 16 to9 around a minute long. Include charts, graphs, and other visuals. Use this lowfi track as the background audio. And then here I just copied and pasted the documentation of how to use Gemini TTS. So basically I just went to Google's AI studio. I selected this add voice over. I selected a speaker and then at the top here I just clicked on get code and copied this code and pasted it into my prompt.
And then at the bottom here I also gave it my API key which I will delete before I publish the video. And that's pretty much it. Let's press run. Now again for my prompt I didn't even tell it what to use to create the video. So, it needs to decide for itself. It decided to use hyperframes, which is an open-source motion graphics generator. So, NX proceeded to write the script, generate the voice over, compose the scenes with charts, etc.
Initially, it hit some errors, so it failed to generate the voice over, but it's really good at automatically troubleshooting and fixing the issues. So, after a few retries, it was able to successfully generate the voice over. And then, next, it proceeded to generate the video. And that's pretty much it. I didn't need to prompt it further. So, here is the final video. >> Five companies, one record-breaking quarter. Here's how big tech's latest [music] earnings stack up.
Amazon crossed $200 billion in quarterly sales for the first time. Alphabet hit 120 billion, Apple 109, Nvidia 96, and Meta 61. But growth tells a different story. Nvidia [music] more than doubled its revenue year-over-year. Meta grew 28%, [music] Alphabet 24, Amazon 20, and Apple 16. Nvidia turned nearly [music] 60 billion in profit at a 75% gross margin. Apple earned 30 billion. Meta's profit fell 14% as AI costs piled up.
The AI engines are roaring. Google Cloud surged 82%. [music] AWS accelerated to 37%, its fastest pace [music] in over four years. The price of that race staggering. Amazon [music] plans 220 billion in capital spending this year. Alphabet 200 billion. Meta up to 145 [music] billion. Looking ahead, Nvidia guides to 108 billion next quarter. [music] The AI buildout is at full steam and Wall Street is watching who gets paid for [music] it.
I would say that it has very good design capabilities. I tried a similar prompt with other Frontier models like GLM 5.3 and GPT, but I do have to say the design and layout of the video from Fable 5.1 do look a bit better. All right, next let me pull up the usage stats for you. So, here is the breakdown. If you want to create the best looking videos with AI, definitely check out Higsfield, the sponsor of this video. They just released Seed Dance 2.5 with 1080p resolution.
Now, Seed Dance 2.5 is hands down the best video generator out there. It can generate up to 30 seconds in a single pass with multiple shots and a full narrative inside one generation. And you can keep extending the video while maintaining the characters, location, pacing, and overall continuity. But previously, it was just 720p. Now on Higsfield, you get to generate full 1080p resolution with Seed Dance 2.5. What's especially impressive is how much control you get.
You can give it up to 50 references at once, including images, video clips, and audio. So you can provide references for your characters, locations, visual style, motion, and even the soundtrack. It also supports multi character consistency, and motion references. If you need to change something afterwards, you can edit specific parts of the video using timestamp level controls without affecting the rest of the generation.
You can also change the background with green screen editing, modify the camera perspective while keeping the characters in action the same, or apply new references to an existing clip instead of just generating clips and hoping for the best. Seed Dance 2.5 gives you a much more direct way to actually control and edit your AI videos. Check out Higsfield and Seed Dance 2.5 1080p using the link in the description below.
Now, these Frontier LLMs cannot generate images or videos by themselves, but you could use them as the director. They can do all the planning and prompting for generative AI models. So, here's a test on its ability to do exactly that. Your job is to make a 30-second commercial 16-9 about this product. I just gave it a link to a random matcha product on Amazon. You may use the images and product specs on that page for reference.
Use Higsfield, MCP, or CLI to generate the content using whichever image or video generator you want. Now, Higsfield is a platform that lets you use a ton of image and video generators. And you can also use their MCP or CLI to control all these models with an agent like Claude or GPT. So, that's what I'm doing here. And then, if necessary, you may generate multiple clips and stitch them together for the 30-second commercial.
Again, I'm going to set it to Fable 5.1 Ultra Code and then press run. All right, here's what I got. It was able to search the web and figure out how this Higsfield CLI even works. I didn't need to give it any documentation. It's able to successfully link my account and then start generating the videos. So, in the end, it generated six clips and then it also decided to use a music model to generate the background music, which is strange, but let's see what it does afterwards.
The music and the voice over generations are ready. So now it is assembling everything together and then afterwards here is the final video. >> Every morning starts somewhere. Ours [music] begins in the tea fields of Kagoshima, Japan. 100% organic matcha. Shade grow and stone ground. [music] Whisk it, pour it, feel it. 36 mg of natural caffeine [music] for steady energy with no crash. more than a dollar matcher. Better boost better noise. >> So, as you can hear, the audio is pretty bad.
There's a lot of overlap. It probably did not detect that the video generations already had audio, so that kind of conflicted with the background music and the voice over. But aside from the audio being all mixed up, the video is actually pretty good. The commercial does use the product specs and the image references from this page. And then for your reference, here's the result from GPT 5.6 Soul with the same prompt. >> Meet Better Boost from Tenzo.
Authentic Japanese matcha, handh harvested and traditionally ground for a smooth everyday cup. USDA organic with 36 mg of natural caffeine for steady energy. Whisk half a teaspoon hot or cold in a latte, a smoothie, or on its own. Premium matcha, less than a dollar a serving. Make every day better. Tenzo. >> Next, let's also test how good it is at music composition. So, here's my prompt. Your job is to compose an amazing Europ EDM song.
Compose the song using any of the VST plugins or samples from my Waveform DAW. You can decide which instruments to use. Be sure to add variations and effects like risers, epic drops, and other elements that make audio files weak to their knees. also include panning FX automation and make sure everything is mixed and mastered properly. Note that I didn't even give it any documentation or reference paths. It needs to figure out where my waveform DAW is located and how to operate it.
Let's click run. So, after some digging, it was able to find the waveform DAW and figure out how to use it. And then it proceeded to write out the notes for each track and compose the song. But then I maxed out the freaking 5 hour limit again. So I had to wait 4 hours and then write try again. In fact, I wrote this too early before the 5-hour limit was reset. So I had to wait a few more minutes and then write try again.
And then afterwards, it proceeded to generate this song. And here is the final result. [music] >> [music] >> Heat. Heat. [music] [music] >> [music] [music] >> Heat up Heat >> [music] [music] >> up [music] here. [music] All >> [music] [music] >> right, that was just part of the song. If you're interested in hearing the full song, I'll link to it in the description below. The chords and structure are actually quite similar to what I got from GLM 5.3.
It added a ton of different instruments like chords, leads, plucks, arpeggio, bass, subbase, pads, keys, guitars, etc. It even added a nice touch at the end there where it kind of transposed the chorus. It's also able to like attempt to mix and master all the tracks. So, you can see it's applying a ton of like equalizers to each of these tracks. There's also different volume and panning going on. Plus, for the master track, it's also applying some standard mastering practices like this equalizer and then a compressor and a limiter.
However, this doesn't really sound very clean or mastered nicely. There's a lot of fuzz or some really high frequencies. Definitely still a long way to go before it matches the quality of a professionally mastered song from a human. Let me know in the comments what you think of this. And if you're interested in hearing the generation from GLM 5.3, I'll also link to it in the description below. And then here are the usage stats for this session.
All right, it's time for your favorite test, finding the frog. So I'm going to upload this image and then I'm going to write, is there any animal or animals in this image? If so, identify and circle it. Let's press run. And unfortunately, here it could not find any animal. So here it says it could not confirm an animal in this image with confidence. There's one area in the lower left center where there is a dark rippled elongated shape which could resemble a copperhead snake which is not correct.
In fact, the frog is also not located in this area. So, unfortunately, this was a fail. But then again, none of the other Frontier models could get this prompt correct. It looks like the frog is safe for now. Now, those were some tests using Clawude Code, but next I'm also going to show you some tests using the regular online chat interface. They claim that Fable 5.1 is worldclass in terms of agentic scientific research.
So let's give it some medical research prompts. For my first prompt, let's try this. Describe the molecular drivers of chronic myogynous leukemia. Discuss targeted therapy evolution, etc., etc. Compare resistance mechanisms and survival outcomes from recent longitudinal studies. Include relevant tables and visualizations. A pretty neutral prompt trying to understand, you know, this type of cancer. I'm not asking it to engineer a virus or anything.
Let me set this to extra and then press run. Now, here's the ridiculous thing about using Fable. At the bottom here, note that it said this switched to Opus 5. I can press edit and retry with Fable 5.1, but it's still just going to revert back to Opus 5. So, I'm not actually using Fable 5.1 for this prompt. It's forcing me to use a dumber Opus 5. Or let's try another example here. I'm just going to ask it to analyze the mechanisms of amaloid, beta, and towel propagation in Alzheimer's disease, etc., etc.
Summarize stuff from recent phase 3 trials. And again, I'm going to set it to extra. And then press run. Unfortunately, even for this prompt, again, it just reverted back to a dumber Opus 5. It's not actually letting me use Fable 5.1. Even though, again, this is quite a neutral prompt. I just want to understand the mechanisms of Alzheimer's disease. I'm not getting it to engineer a freaking pandemic, but still, it's not going to let me use Fable 5.1.
I'll talk more about this limitation in a second. All right, final test. I don't expect this to work, but let me also upload this image of six different brain scans. Each of these scans contain a different type of brain tumor. I'm going to ask Fable 5.1 to identify the types of tumors in each of the six images, if any. Let's press run. Now, interestingly here, it did not revert to Opus 5, which is so strange. But anyways, here is its answer for the top left.
It's giving me all these different options. So, I'm going to assume just its first choice is the highest probability, which is wrong. This is not maccoinoma. And then for the top middle, it said lepto menial carcinomasis, which is wrong. And then for the top right, it says meniomas, which is also wrong. For the bottom left, it said there's no tumor evident, which is wrong. For the bottom middle, it said calcified cranoparingioma, which is also wrong.
And then for the bottom right, it said no tumor, which is also wrong. So, it got six out of six wrong. But note that this is a really hard prompt. So far, the best performers are GLM 5.3 and Kim K3, which only got one out of six correct. All right, finally, let's test it on its creativity and ideation. There's no right answer to this, but here I just want to see how good its ideas are. Give me five simple tech startup ideas that don't exist yet and have the highest chance of making 10 million ARR within a year.
Let's press run. All right, so here is its response. I'm not going to read out the whole thing, but the first idea is an AI agent expense auditor where you can sell it to CFO or engineering ops. The second idea is a compliance ready AI output archive and this is for regulated firms. They must retain and prove what their AI systems told customers or employees. So they need immutable logging plus retrieval plus audit reports.
Another idea is voice agent QA and red teaming which you would sell to call centers. The fourth idea is automated vendor security questionnaire response for small and medium-siz businesses. And then finally we have agentto agent payments. Again there's no correct answer to this. Let me know what you think of the quality of these ideas. All right. So, that sums up my series of diverse tests on Claude Fable 5.1 so you can get a sense of what it can and cannot do.
Next, let's go over its specs, performance, limitations, and where you can use it. So, here is their official release page. And if you scroll down a bit, here is a table showing its performance against the previous Fable 5 as well as Opus 5 and the best GPT 5.6 Soul. You can see in particular for this terminal bench science which tests its ability on agentic scientific research. Fable 5.1 completely destroys its competition scoring 52.6% which is like more than double most of the other models.
However, this means absolutely nothing to me because like I showed you with my medical prompts, it just constantly reverts back to a dumber Opus 5. It doesn't actually allow me to use Fable 5.1. So, it's pretty much unusable for most like medical research or biology related tasks. Next, in terms of agentic coding, again, it's state-of-the-art, even significantly beating GPT 5.6 Soul. Same with these knowledge work and agentic use benchmarks, as well as humanity's last exam, which tests the AI models knowledge on some pretty obscure domains.
Same with business workflows and agentic coding. Here it says that Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads, which is outright lying. I'll talk more about that in a second. Now, next, here are some of its impressive achievements that they listed on this page. Now, using Claude Mythos 5.1, which is basically Fable 5.1, but without guardrails, but unfortunately, you plebs are not allowed to use this.
Here, it was able to design proteins that actually worked in the lab. So it generated protein binders. They tested this on three targets and the best designs from mythos 5.1 turned out to bind 10 times more strongly than the previous best proteins which is pretty crazy. And also nearly 50% of its designed proteins work whereas the regular success rate of protein design today is only like 10 to 15%. But again you plebs cannot have access to this model which can design proteins.
Another cool achievement is that it created a much better map of planet Venus. Starting from decades old NASA radar data, it trained a neural network and produced a highresolution elevation map showing about 1/3 of Venus, improving the detail from roughly 10 to 20 km down to only 2 to 3 km in resolution. It also improved the height accuracy by up to 25%. They also got Mythos 5 to optimize scientific AI models dramatically.
So, it rewrote parts of seven biology and genomics models and made them as much as 2.5 times faster, cutting GPU costs by 30 to 60%. Enthropic says that this kind of optimization would take engineers weeks to do. But here, Mythos 5.1 was able to just complete this in days. But again, keep in mind, you can't even use this model. So, those were the highlights from their official release page. Next, if you look at this independent leaderboard by artificial analysis, according to their intelligence index, then Claude Fable 5.1 does score the highest at 66 points, whereas the best GPT 5.6 Soul is only 61 points.
However, the thing I don't like about this leaderboard is that there aren't any confidence intervals. So, I mean, the Frontier models only differ by like around five points. I'm not sure if this is actually statistically different, but anyways, according to this intelligence index, then Fable 5.1 with fallback is indeed the most performant model in the world. However, if you look at the cost of this, here is where things start to fall apart.
Contrary to anthropics claim that this is a quarter cheaper than Fable 5, you can see that artificial analysis found that it's actually way more expensive. So, you can see it costs $3.69 69 cents per task whereas Fable 5 is only 3.14. And if you compare this to the other Frontier models like GPT 5.6 or Quen 3.8 or Kimikate 3, GLM 5.3, this is like four to eight times more expensive. Definitely not the most costefficient option out there.
In fact, here's a crazy chart showing how much it cost the artificial analysis team to actually run their evaluation. So, it cost them a whopping $8,500 to evaluate Fable 5.1, whereas for Fable 5, this is like way lower and even lower for the other competitor models. So, I mean, this claim from Enthropic that Fable 5.1 is cheaper is just blatantly lying. Even for me, I found that Fable 5.1 not only costs more, but it also drains my usage limit way faster than Fable 5.
Fable 5.1 is also slow as hell. So here's the latency or basically the time to first answer token. And as you can see, Cloud Fable is by far the slowest. Basically almost three times as slow as GPT 5.6 Soul. And if you compare this to like GLM 5.3, then Fable 5.1 is like 10 times slower. Again, here's another chart showing the endto-end response time. Fable 5.1 is the slowest model out there. And if you look at this omniscience hallucination rate, then again, Fable 5.1 is not the best performer.
It actually hallucinates more than even Fable 5 as well as some other open source models like Kim K3 and GLM 5.3. You can see that Fable 5.1 hallucinates like more than twice as much as GLM 5.3 Max. If you look at this other leaderboard by Arena where people can blind test different AI models side by side, then it doesn't look like they've added Fable 5.1 to most of these leaderboards, but they have added it to webdev.
And as you can see, Fable 5.1 is by far the best performer. significantly beating the second best model, Quinn 3.8 Max, which was just released yesterday. Here's another leaderboard called Livebench by Abacus AI. And as you can see, Clot Fable 5.1 is indeed ranked number one on this leaderboard as well. In fact, it on average scores the highest across all these different categories, including reasoning, coding, agentic coding, mathematics, language, etc. on Deep Suite, which test the models ability on some realistic agentic software engineering tasks.
They haven't added Cloud Fable 5.1 here, but shockingly, we have a new model, Gemini 3.8 Flash, that was released today, which is now ranked number one. That's pretty crazy. I'll talk more about Gemini 3.8 in a later video. Now, while on average across multiple independent leaderboards, Fable 5.1 is ranked number one, there are a ton of red flags you need to be aware of. In addition to being the most expensive, you should be aware that all the responses generated by Claude models have a hidden text watermark, which means that others could take your response and plug it through a verifier to confirm if that response was actually written by Claude.
Now, for most cases, it wouldn't really matter, but if for whatever reason you don't want people to know that you used Claude to generate your response, then this might matter to you. Another red flag is, as I've shown you throughout in the video, it eats up my 5hour limit so quickly. Like after one or two prompts, it already maxes out my 5-hour limit, and I need to wait like several hours before I can continue. And that's why, you know, this video took over a day to release.
Like, it took roughly 15 hours to complete all the tests in this video. And the majority of the time is just waiting for the damn 5hour limit to reset. So, this was quite a painful user experience for me. Another red flag is that Anthropic has done some really deceptive lying and marketing. In fact, they're getting sued for this right now. So, here is the lawsuit. Basically, on their site, they specify that if you pay for the Max plans, you get five times or 20 times more usage.
Note that on their pricing page, this is exactly what it says, five times or 20 times more usage than Pro. There's no asterisk or footnote anywhere explaining what this means. But it turns out that this is just five times or 20 times more of the 5h hour limit, not the total usage. In fact, they calculated that the total usage is only like 3.5 times if you use the 5x plan or six times if you use the 20x plan. So, this is extremely deceptive.
Here's another disgusting thing from the Claude team here. They deceptively posted this. Starting September 14th, we are permanently raising standard weekly limits by 25% for all these plants. So, this sounds great, right? They're raising the limits, but as this community note pointed out, this is actually a net 17% reduction compared to the active 50% promotional limit. So, instead of getting 25% more, you're actually getting 17% less.
Again, incredibly deceptive. And here on their official page here, they blatantly lied that Fable 5.1 costs less than Fable 5. But as you can see from Artificial Analysis, this is far from true. By the way, because of this lawsuit, because they were caught lying, if you're looking to cancel your subscription, don't just cancel, but you can also try talking to their customer support chatbot and requesting a refund. Get your money back.
And finally, here's another disgusting thing about their marketing. If you want to actually try out Claude Fable 5.1, you can do so using their API, but if you want to use it through a subscription plan, you have to use the max plan, which is at least $100 per month. The pro plan actually does not offer access to Fable 5.1. Again, they don't even mention this over here. Now, this is turning into a rant, but if you look at all these independent leaderboards, then on average, Claude Fable 5.1 is the most intelligent model out there right now.
So, if you need an AI to do some incredibly challenging coding or reasoning stuff, as long as it's not related to cyber security or biology or chemistry, and if the other frontier models are not able to handle the task, then as a last resort, it might be worth trying Fable 5.1 and seeing if it can figure it out. But in most cases, at least for myself, just due to the horrible marketing tactics of Anthropic, plus they've had a long history of gatekeeping and rugpulling their users and fear-mongering from their CEO.
I personally would not use Fable or subscribe to a Claude plan. Anyway, that sums up my review of Cloud Fable 5.1. Let me know in the comments what you think of this. As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week. I can't possibly cover everything on my YouTube channel.
So, to really stay uptodate with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below. Thanks for watching, and I'll see you in the next one.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.