Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Search · @theAIsearch
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Search's most watched videos.
Most replayed moment at 10:28
3.1x that video's typical replay level
and then fix any errors that it sees. And so afterwards, it rendered the video successfully and that's pretty much it. In just one prompt, I didn't even need to prompt it further. Here's our final result. >> Four companies, one quarter, and a half-trillion-dollar bet on artificial
Said at 10:22
Most replayed moment at 15:21
4.1x that video's typical replay level
and deep agent for only $10 a month. This is way cheaper than if you paid for each tool separately. Definitely check out chat.llm that comes with deep agent in the description below. You can think of the residual connections not as a simple pipe carrying the signal forward,
Said at 15:15
The graph counts replays. It does not show where viewers stopped watching.
Words
4,500
Runtime
25:55
Speaking pace
174wpm
Reading time
19min
174 words per minute, between the 160 25th percentile and the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Claude Opus 5.5 is out, and this might be the best model in the world. In this video, we're going to go over what it can and cannot do, plus we're going to go over its specs, pricing, and benchmarks. Let's jump right in with some demos. Now, you could use Opus 5.5 using the online chat interface, but as with most of the Frontier models out there, this is designed for a gentic coding and really long horizon tasks. You can just give it a goal
87 words, the words spoken in the first 30 seconds at 174 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
Claude Opus 5.5 is out, and this might be the best model in the world. In this video, we're going to go over what it can and cannot do, plus we're going to go over its specs, pricing, and benchmarks. Let's jump right in with some demos. Now, you could use Opus 5.5 using the online chat interface, but as with most of the Frontier models out there, this is designed for a gentic coding and really long horizon tasks. You can just give it a goal and it can work across many different platforms and call different tools autonomously for hours until it achieves your specified goal.
And the best way to do that is to actually use Claude through a harness. And the best one to use for Claude is Claude Code. This allows you to work on multiple projects at once. Plus, you can also link to your local files and folders. So, that's where I'm going to show you most of my tests. For the first test, let's try if it can bypass different recapture puzzles. I'm going to tell it to open a browser and then go to this page and complete all the tests.
Show your clicking and typing so I can see your work. I'm going to select Opus 5.5. Plus, I'm going to set this all the way to max. Now, I could set this to Ultra Code, but this is just going to spin up hundreds of agents and drain my usage fast. So, I'm just going to set it to max so it doesn't spin up concurrent agents, but it should be the same level of intelligence. Let's press run. All right. So here you can see it spinning up a browser beside the chat.
Now the first recapture is pretty easy. For the second one, it has to select all the squares that contain the stop sign and it also was able to do this really quickly, like in just a few seconds. And then here's the third test. It needs to input the text. It's also able to figure this out really quickly. In fact, it typed so fast that I cannot even see it. Anyways, here's the next recapture test. So it only selected two cells here, which is wrong, and now it's figuring out what to do next.
But it was able to select the other vegetables in the second pass and then get through to the next level. And then here is where it needs to rotate these cells so that the photo is coherent. And it was able to do this very well. Next, it needs to play tic-tac-toe against the computer. So for the first round, there's no way for it to win. It was a tie. So here's the second round. And for the second round, it was able to win at tic-tac-toe.
All right. Next, we have this word search puzzle. It's actually pretty easy. So, it was able to find the specified words. Next, it needs to select the squares with the stop sign. But here's the trick. Once it clicks on a cell, it actually becomes a nested grid where the cells become smaller. So, it figured that out along the way, and then it was able to successfully select all the smaller cells that contain this stop sign.
All right. Next, it needs to whack a mole. This is super tricky because these moles change location over time. So, I'm not sure how it can actually beat this stage because Opus does take a long time to think before it executes. So, by the time it finishes thinking, the mole will be gone. So, let's see what it does. But after a few minutes, it was able to figure out that it should just wait at a cell and then click it the moment the mole appears.
And in just a few minutes, it was able to complete this test. Next, here's a where's Waldo test. Let's see how quickly it's able to find Waldo. Oh my god, that was way too quick. So, it only took like less than a minute for it to find Waldo. Here's the next level. It needs to identify the Chihuahua. Also, super easy for it. And then here are some additional tests that it had to do. Let me fast forward to this test where it keeps getting stuck.
So, it needs to navigate the car to the parking spot, but it keeps crashing into the corner over here or it keeps crashing into other cars. So, it spent a really long time on this test, like over 10 minutes, but eventually it was able to park the car. However, for like a real recapture, 10 minutes would be way too long. So, I would say this is a fail. All right, so those are some tests of its real-time adaptation. All right, next.
Let's test its physics understanding and ray tracing abilities. So, I'm going to write make a ray tracing 3D simulation of a bullet piercing through a water balloon, making the balloon pop and the water bursting, deforming from the original shape, and eventually falling to the ground. Allow me to freeze the animation at any point to view the scene from any angle. Add different sliders for settings like bullet speed, balloon size, etc.
And here's the most important part. Do not use 3JS or any external web libraries. You need to code this up from scratch. And also, here is a really useful prompt hack to enhance the quality of its output. So, for each iteration, use a separate critic agent to take screenshots of the scene to evaluate its realism and accuracy. It should provide a score from 0 to 10. 4 to 7 means it's usable, but would not pass. Eight or above means it looks very convincing.
Loop until the critic agent gives a score of eight or higher. Below that, the builder gets the ranked issue list and tries again. And do this for up to three rounds. Again, I'm going to set it to max and then press run. All right, here's what I got. If I scroll down a bit, you can see that after it finishes coding everything, it's using the critic agent to give it a score. For round one, it only gave it four, so it had to loop again.
For round two, it gave it 4.8. Still pretty bad. So, it had to loop for a third time. For the third time, it actually gave it 4.6, which is even worse. But I did set the maximum number of loops to three times. So, it ended there. Now, I tried opening this up and the page was frozen as if something was stuck. So, I wrote, "Make sure this is efficient enough to run on a regular web browser." So, it continued refining and coding it some more.
And afterwards, here's our final result. All right. So, here is our simulation. I can adjust the slider at any point in time at the bottom here. So, here you can see the bullet piercing through the water balloon and then showing the water dropping to the ground. Very nice. So, let me adjust the slider here. I can drag this around to view this at different angles. I can also decrease the bullet speed, caliber, vertical offset.
Let's also increase the balloon radius. There's also like latex tear speed, opacity, latex color, and then for physics, I can also adjust the gravity, energy into water body, surface tension, cohesion, viscosity, spray amount, mist amount, etc. It gave me a lot of settings to play with. So, let's try to burst this larger balloon. So, that's what it looks like. And then down here, I can even set it to a different environment.
So, let's try studio, which looks like this. Light elevation, light intensity. So, a ton of different settings that I can play with. Now, for your reference, here is the generation from GPT6 Astra. I actually like the look of Astra's generation a bit better, but let me know in the comments what you think. All right, let me pull up the stats for you. So, this took roughly an hour and a half. And here are the stats for this session.
Let's increase the difficulty even more. I'm going to get it to go to this site, which is just a virtual online piano. And I can click on these keys here to play the notes. Set visible keys to max. So sometimes over here in settings, if it's not set to max, then you get fewer keys. So I need to set it to the max number of keys. And then compose a fast-paced piano solo in the style of Shopan around a minute long and play it by pressing the keys on this online keyboard.
So the challenge here is first it needs to compose this song, but it doesn't just play the song in a DAW. It needs to actually play out the song live by pressing the keys sequentially. And the timing of each key press has to be correct in order for this to sound good. So there's also a dimension of time here. Let's see what it does. All right. So first it's spending a few minutes trying to figure out how to actually press the keys on the keyboard, how to record its performance, and then after learning how this thing [music] works, here is its final performance. >> [music] [music] >> Holy smokes.
That was insane. That sounded even better than what I got from GPT6 Astra. Now, this did think for a super long time, so it's around like 30 minutes. And let me expand the usage stats for you. Next, let's see how good it is at creating math explainer videos. So, here's my prompt. Create a motion graphics explainer video. It should be minimalist, white on black, around a minute long. The topic is how we first figured out the Earth's circumference.
It should be easy to understand but go in depth on the mathematical concepts behind it. Now Claude cannot generate a voice over itself. So I told it to use Gemini TTS. I pasted some documentation on how to run it. And that's pretty much it. Let's press run. All right. So this also worked for around half an hour. And here's my result. Around 240 BC, a librarian in Alexandria measured the whole earth with a stick, a shadow, and geometry.
Erosines heard that insin to the south the noon sun on the summer solstice shone straight down a well no shadow but in Alexandria at that same moment a vertical stick did cast a shadow the tangent of the angle equals shadow over height about 7.2° 2°. Here's the key. The sun is so far away that its rays arrive parallel. Extend both sticks to Earth's center. A line crossing parallel lines makes equal alternate angles. So, the angle at the center is also 7.2°.
And 7.2 over 360 is 150th of a circle. So, the 5,000 stadia between the cities is 150th of the circumference. Time 50, 250,000 stadia, about 40,000 km. The true value, 40,075. One angle, one distance, and he measured the world. >> Pretty good, I must say. It's slightly better than what I got from GPT6 Astra. And then here are the usage stats for this session. If you want to turn Claude Opus 5.5 into a full creative production team, definitely check out Higsfield, the sponsor of this video.
Most people are still using Claude inside a chat box, but with Higsfield MCP, Claude can actually generate the images, videos, and other creative content. The combination is pretty powerful. Opus 5.5 acts like the creative director, while Higsfield handles the actual production. You can give it a single creative brief, and Opus can plan the entire workflow, write scripts and shot lists, generate the individual scenes through Higsfield, review the results, and even improve its own prompts, and regenerate anything that needs to be fixed.
For example, you can ask it to create a cinematic product ad from just one product image. And Claude can come up with the concept, break it down scene by scene, and directly call Higsfield's image and video models to generate everything without you having to copy prompts between different tools. Because Opus 5.5 is designed for longunning agentic tasks and has a massive context window, it can keep track of your creative brief, previous generations, feedback, and visual direction throughout an entire project.
Higsfield MCP gives Claude access to models like Seed Dance 2.5, GPT image 2.5, and the rest of Higsfield's creative model stack. So whether you want to make short films, music videos, product campaigns, content series, or anything else, you can just go from a simple idea to finished visuals through one continuous workflow. Try Higsfield today using the link in the description below. All right, next. Let's see how good it is at creating 3D models in Blender and coding up a full real estate virtual tour.
So, I'm just going to give it this link to a random Airbnb listing. Let me show you some photos first. Here's the living room. There are some sofas and then the dinner table over here. There's a pool outside. So, here's another view of the living room. There's this weird wooden object over here. Here's another view. Note that the pool also has this slanted wall. And then here's the bedroom. Here's the washroom. We don't know exactly where the washroom is relative to all the other rooms.
And then here's the exterior. All right. So, some additional photos for your reference. I'm just going to give it that page and then get it to reconstruct the entire building and interior using Blender MCP at this address. So what I did in Blender is I already installed this Blender MCP and I just need to click on connect to MCP server to have it running on this local address so that any AI model can now control it and then afterwards program the camera path to fly through the property like a virtual tour then render the scene.
All right, now this worked for a ridiculously long time like over 4 hours. But eventually it did give me a really nice render of everything. In fact, it even coded up the environment around this property, which is crazy. Afterwards, here's what we got. In fact, let me open up this Blender project first so you can see how ridiculously detailed this is. So, here's the whole freaking property. And if we zoom in a bit. In fact, let me turn on the surface of this as well, so you can see it.
So, here's the pool. And if we zoom in, here's the living room. The textures are pretty decent. So, here's the pool, here's the bedroom, and here's the washroom. Everything looks pretty decent. And then here is the final render. [music] >> [music] [music] >> It's not entirely faithful to the original listing, especially the pillows, but this does look slightly better than what I got from GPT6 Astra. Pretty impressive.
And then for your reference, here are the usage stats for this session. Now, coding up a model and creating an animation in Blender is relatively easy. Let's see if it can code up an entire video game, including animating the characters and creating the entire environment in Unreal Engine. So, here's my prompt. Create a procedural 3D game in Unreal Engine. The user controls this character in thirdp person point of view.
So, I just linked to this character on Sketch Fab. And then I wrote, "She's able to sprint quickly and jump very high, leaping across rooftops like a ninja. You need to add appropriate animations for the character. You can look for relevant animations in Miximo and map them onto your character. So Miximo is this free site which contains a ton of different animations which you can add to your 3D models. And then the environment should be a procedural ancient Chinese imperial setting.
Generate all assets in Blender and then render them procedurally in Unreal Engine. Attached are some photos of the environment for your rough reference. So, I gave it these three photos and then I wrote make it as detailed, realistic, and grandiose as possible, like a AAA video game. And again, I used this prompt hack where it needs to spin up a separate critic agent that takes its own screenshots from separate viewpoints and scores the game design and aesthetics from 0 to 10.
Over 8.5 is AAA quality. If it scores below that, the builder gets the ranked issue list and tries again up to four rounds. Let's press run. All right, so this took around 4 hours. It just kept working and working. It was able to take and capture some screenshots itself like this. It was able to find issues and fix them autonomously. Here it's taking some additional screenshots. And then afterwards, here's our result.
So, let me start running around here. I can jump across rooftops and sprint. Everything looks very beautifully designed. The environment looks great. There are a lot of subtle details. This actually looks better than what I got from GPT6 Astra. the lanterns, the cauldrons, the floor and the stairs, the design of the walls, everything just looks a lot more detailed and professional compared to GPT6 Astra, which if you looked closely, still contains a ton of flaws.
So, you know, Opus 5.5 is pretty good at creating 3D video games like this. Next, let's see how good it is at creating music and operating my DAW. So, here's my prompt. Your job is to compose an amazing EDM song around a minute long. Compose the song using any of the VSST plugins or samples from my DAW. You can decide which instruments to use. Be sure to add variations and effects and other ear candy elements that make audio files drool and weak to their knees.
Make sure everything is mixed and mastered professionally. Let's press run. All right, so this worked for an hour and a half. The first render actually sounded horrible, so I wrote there's too much reverb for some tracks. Make it sound tighter and cleaner. So it worked for a bit more. And here's our result. Heat. Heat. [music] [music] Heat. Heat. [music] >> [music] [music] >> So, that was what we got. It's not too bad, but still not close to a professionally mastered track from a human.
Let me know in the comments below what you think of this track. And then for your reference, here are the usage stats for that session. Now, in addition to clawed code, let me also show you some examples using just the online chat interface. In fact, let's do your favorite test. Finding the frog. I'm going to feed it this image and then ask it, is there any animal or animals in this image? If so, identify and circle it.
And then I'm going to select opus max and then press run. All right. So, it thought for a few minutes. It tried to zoom in on like multiple parts of the image. And then here it's confirming that there's a lizard here which is completely wrong. And then here's its answer. It says there's a lizard and it's in the left center of the photo circled here which is completely wrong. So unfortunately even the current best model in the world Opus 5.5 is not able to pass the frog test.
The frog is safe for now. All right. Next, let's see if it can identify cancer. So I'm going to upload this photo of six different brain scans. Each scan contains a different type of brain tumor. I'm going to upload this image and write identify the types of tumors in each of the six images if any. All right, for top left, it wrote menioma, which is correct. For the top middle, this is wrong. For top right, this is also wrong.
And then for the bottom left, it says it could not see a clear tumor, which is wrong. For bottom middle, it wrote cranopioma, which is wrong. For bottom right, it wrote this, which is also wrong. So, it got one out of six correct, which is actually state-of-the-art. So the other models like GPT6 Astra and also GLM 5.3 and I believe Kim Free were also able to get only one out of six correct. This is the highest score we have so far.
All right, next let's see if it can do deep research. So I'm going to get it to analyze the pathophysiology of atherosclerosis include relevant tables and visualizations. And here's my result. It's able to give me a nicel looking flowchart plus a table with some very concise information. Here's the next table. strategy comparison, major outcome trials. I'm just going to scroll all the way down so you can see its full results.
Now, there's no right answer to this. All the Frontier models can also do deep research. So, this is more of a subjective preference. Each model kind of has a different vibe. I would say the Claude models are sometimes too concise and they sometimes leave out information. I'm not sure whether that's intentional or not. So far, the best deep research tool I've used is actually Kim K3 followed by GPT6. Finally, let's also test its creativity and its ability to think of new ideas.
Propose the three highest impact ideas that don't exist yet that use current tech and AI to meaningfully improve the lives of animals in factory farms. Hard constraints, meaningful improvement in quality of life, can scale across a huge number of animals. Low effort and operational cost. So, adding this would be a no-brainer for operators. Let's press run. All right, here's the results. Number one is chick distress call autopilot.
So with a few microphones, we would get AI to track the rate and location of chick distress calls, which would adjust the heat, ventilation, and other environmental settings. And then number two is drinker line autopilot, where we have a ceiling camera that judges drinking posture and spots wet patches, which will then be used to set the drinker height and water pressure. This matters because a wet floor is a major cause of this footpad dermatitis.
So there should be some monitoring and fixing of this. And then idea number three is a phone app plus a cheap smart relay on the existing air raator. So this would be for fish and shrimp farms. Again, this is quite subjective, but let me know in the comments below what you think of these ideas. Do you think they are novel and creative? And do these ideas actually meet my criteria? So that sums up my tests on Opus 5.5.
It's incredibly good at design, creating videos, making 3D models, video games, and even playing the piano. Next, let's go over the performance and specs of this. First of all, here are their self-reported benchmarks. You can see across the board, at least for these agentic coding benchmarks, as well as knowledge work, scientific research, computer use. It even beats their previous best model, Fable 5.1, as well as GPT6 Astra.
It's also much cheaper and more efficient. If you look at this benchmark for a Gentic terminal coding, the X-axis is the cost per attempt, and you can see that in general, it costs less than GPT6 Astro while achieving a higher score. If you look at this artificial analysis intelligence index, then you can see that the latest Opus 5.5 is ranked number one. Now, as with all the clawed models, this is incredibly expensive and I think overpriced.
It costs around $6 per task. Keep in mind that this is already cheaper than Fable 5.1, which is even more expensive and restrictive. But then again, this is still way higher than GPT6 Astra. If you look at this omniscience hallucination rate, then you can see that at least among the other clawed models, Opus 5.5 does have the lowest amounts of hallucination, but it still hallucinates more than GPT6 as well as Grock and Muse.
If you look at another leaderboard called Livebench, then you can see that they ranked Opus 5.5 in second place, slightly below Fable 5.1. Now if you look at mazebench which tests how good an AI model is at navigating mazes then you can see that clot opus 5.5 is still way behind GPT6 Astra. If you look at Val's index which is how good a model is across various knowledge work tasks including finance coding and legal then you can see that opus 5.5 is ranked number one whereas GPT6 Astra is in fourth place.
If you look at kernelbench, which tests how good an AI is at generating really complex, low-level code for GPUs, you can see that Opus 5.5 is ranked number one, much better than GPT6 Astra, which is all the way down here. For other leaderboards like ARC AGI 3, as well as Agents Last Exam and Deep SWE, it doesn't look like they've added Opus 5.5 yet. Finally, it's also important to mention the guardrails for this. So, as with the previous frontier clawed models, Opus 5.5 also has similar guardrails to Fable 5.1.
So, if it thinks you're asking about some dangerous cyber security or biology questions, or if you're trying to distill from the model, it will automatically fall back to a dumber model. Finally, let's go over where you can use it. So, Cloud Opus 5.5 is available on all paid plans and via API. If you're on the free plan, you won't have access to Opus 5.5. That sums up my review of Claude Opus 5.5. At least for certain tasks, this does seem to be the best model in the world.
And interestingly, they've even made it cheaper than Claude Fable, which was very restrictive. Let me know in the comments below what you think of this. As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week. I can't possibly cover everything on my YouTube channel.
So, to really stay uptodate with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below. Thanks for watching and I'll see you in the next one.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 361 |
| Average words per sentence | 12.5 |
| Longest sentence | 61 words |
| Questions asked | 3 |
| Sentences containing a number | 58 |
Most used terms
Filler phrases
39 in total: like 23 · actually 14 · kind of 1 · you know 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.