Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Search · @theAIsearch
Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
17:586.9x the video's typical replay level
time for my famous frog test which no AI model has gotten correct so far. So, I'm going to upload this image, and yes, there is a frog hidden somewhere in this image. If you want to look for the frog, you can pause the video here and try to find it. And then, I'm going to ask it to find and circle the frog in this
Said at 17:50
Most replayed moment #2
14:423.2x the video's typical replay level
here's our results. So, we have a piano, synth pluck, strings, drums, and bass. Let's hear its award-winning composition. >> [music]
Said at 14:35
Most replayed moment #3
25:512.7x the video's typical replay level
patterns on the fly. And as you can see, GPT-5.5 extra high actually scores really well, like 85%. Finally, let's also look at how often it hallucinates. So, GPT-5.5 extra high is ranked over here, which is not good. So, it
Said at 25:44
The graph counts replays. It does not show where viewers stopped watching.
Words
4,942
Runtime
27:13
Speaking pace
182wpm
Reading time
21min
182 words per minute, just over the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
OpenAI just released their latest model, GPT 5.5, and this is now the best and most performant AI model you can use. This is a huge upgrade if you use it properly. So, in this video, I'm going to show you all of the incredible things you can do with it. Plus, we're going to go over its specs and where to use it, and finally, we're going to compare its performance and benchmarks against other competitor models. Let's jump right in. Thanks to HubSpot for sponsoring this video. First of
91 words, the words spoken in the first 30 seconds at 182 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 396 |
| Average words per sentence | 12.5 |
| Longest sentence | 47 words |
| Questions asked | 7 |
| Sentences containing a number | 54 |
Most used terms
Filler phrases
52 in total: like 31 · actually 11 · I mean 3 · you know 3 · basically 2 · kind of 2.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
OpenAI just released their latest model, GPT 5.5, and this is now the best and most performant AI model you can use. This is a huge upgrade if you use it properly. So, in this video, I'm going to show you all of the incredible things you can do with it. Plus, we're going to go over its specs and where to use it, and finally, we're going to compare its performance and benchmarks against other competitor models. Let's jump right in.
Thanks to HubSpot for sponsoring this video. First of all, note that all the top models out there can already do simple stuff like helping you summarize things, writing an essay, writing emails or social media posts, and other simple stuff. So, I'm not going to test any of these in this video. I'm going to give it way trickier tasks and really push it to its limits so you can see how performant this is. Now, there are several places where you can use GPT 5.5.
Of course, you can use it on ChatGPT, but especially if you're trying to build or vibe code something, then it's highly recommended that you use their Codex app instead, which you can download and try for free. Because in Codex, you can basically get an agent or multiple agents to work on an entire project in a folder. For example, for this video, you can see that I have all these different folders, and each is a separate project, which I'll show you in a second.
And this allows you to build way more complex stuff and also keep iterating on your project. This is much more effective compared to just using the ChatGPT interface. Let's start off with a really hard test already. Build a fully interactive 3D digital twin of Earth that allows users to zoom seamlessly from outer space down to individual city streets. Show a realistic Planet Earth. Use publicly available assets, models, and layers if needed.
Make sure it loads efficiently on a regular web browser. And then for this, I'm going to select GPT 5 extra high, and then press run. Now, this is a thinking model, so it actually thought for quite a while before it gave me this output. And this output is already pretty good. But the 3D buildings were kind of lackluster, so I wrote make the 3D view look better. It should span all buildings. And then afterwards, it was kind of laggy, so I wrote the buildings should be loaded more efficiently for a regular web browser.
And finally, here is what we get. So, here is what the Earth digital twin looks like. It even was able to map the cities onto this globe in 3D. Very nice. Now, if I click on night, this has night lights enabled. Very nice. All right, next, let's click on New York and it automatically pans me to New York. Let's turn off night now, and then let's click on San Francisco, and indeed it zooms into San Francisco. All right, next, let me try to click on this button to dive into street view, and indeed it was able to give me this street view.
Next, let's also click on 3D buildings, and it's able to render all these 3D buildings. How cool is that? Now, if I disable streets, you can see just the 3D buildings like this. Really cool. In just a few prompts, it's able to code up everything that I specified in the prompt. It's able to give me a street view map with 3D buildings. It's also able to give me a planet view with night lights. I can probably add other layers here as well.
This is pretty awesome. So, that's one example of some pretty complex things that you can get it to code up. Next, I'm going to start a new project, and then here, let's develop a ray tracing simulation featuring one sphere, one cube, and one pyramid. The environment is a blue sky is a blue sky with a checkered ground. Add adjustable parameters like position, reflectivity, roughness, translucency, and other material properties of the sphere.
Put everything in a standalone HTML file. Make sure it loads efficiently on a regular web browser. This is a key phrase that I like to use to make it run smoothly. Let's press generate. All right, so here's what I got at first and it already worked right out of the box, but notice that in the prompt, I only got it to give me the settings for the sphere. So, next what I did was prompted it to add the same sliders to the other shapes as well.
And then afterwards, it seems like the reflections of the other shapes weren't really rendered correctly, so I wrote "Make sure each shape's material properties are also correct when reflected in another shape or seen through a translucent shape." That's pretty much it. And just like three prompts, here's what we got. So, let me open this index.html file and then I can also click on this button to open it in an external browser.
Indeed, we have these three shapes. Let's adjust the position of the sphere. So, these sliders work. Let's adjust the radius. So, radius also works. And then reflectivity, let's make this super reflective and let's also decrease the roughness. And then for translucency, let's change this all the way to zero, so it's like a really reflective metallic sphere. And as you can see, the reflections are correct. You can clearly see the reflections of the shapes in the sphere.
Now, if I drag translucency all the way to one, you can see that now it turns into a very transparent sphere. Really cool. And then I can also play with this IOR setting, which does this. Very nice. And then the specular setting. This is very subtle, but it adjusts the surface of the sphere. And then let's also change the color to something like green. So, color also works. In fact, if I decrease translucency, then you can see the color a bit better.
Let's set this back to blue. All right, so that's the sphere. Let me also quickly test the sliders for the other shapes. So, position works for the cube. Size also works. Reflectivity also works. Very nice. And translucency also works. And then IOR also works. And if we look through the cube because it's like 90% translucent, you can clearly see the sphere through the cube. And then, finally, let's also test the pyramid really quickly.
So, the position and the size work. Let's also increase the reflectivity. So, that works. Notice that the reflection of the pyramid in the sphere also changes as we change the reflectivity. And then, let's also adjust translucency. So, that works. Very nice. And let's set the color to white. All right. So, right now we have the pyramid and the cube translucent and the sphere is metallic. Finally, let's also adjust the sphere to completely translucent.
And here's what we get. So, it's able to render all of this in just three prompts. Really impressive. All right. Next, I'm going to also show you some examples in ChatGPT. So, if you click on this model drop down, you can click on configure and then select Pro 5.5 here. And then, for the thinking effort, let's set this to extended so you can see the most performant version. All right. Let's see if it can identify cancer.
I'm going to paste in this image and then ask it to describe what this photo is about. If there are any lesions in the photo, circle them. Let's press run. It is correct that this is a montage of axial chest CT slices. Let's download the annotated image and see if it was able to circle the lesions. All right. On the left is GPT's answer and on the right is the correct answer. You can see for slide one, it was not able to circle the lesion correctly, which should be over here.
For slide two, it was able to get it, but this one is pretty obvious. For slide three, it also got it correct. And for slide four, it was also able to get it correct. So, not completely perfect. It was able to get three out of four right. Everyone's learning AI right now, but most people hit the same wall. You can use the tools, you just don't know how to turn that into something that actually makes money. You've probably thought of building an AI app, but figuring out what to create or whether your idea will actually make money can feel overwhelming.
That's why you should check out the 50-plus AI app ideas making millions database by HubSpot. I put it in the description below so you can access it for free. This database gives you access to dozens of real app ideas that are already generating revenue. You can explore thousands of data points to see what's working right now in the market. It lets you filter ideas by industry, platform, and product type. This makes it easy to find opportunities that match your skills, interests, or business goals.
Plus, each entry includes links to full case studies so you can dig deeper and understand how these apps were built and how they actually make money. My favorite part is how practical and easy to use this resource is. Instead of guessing what might work, you get real examples of AI apps that are already succeeding in the market, which makes it much easier to validate your next idea. You don't need to reinvent the wheel here.
You can access the full database for free using the link in the description below. This resource was created by HubSpot, the sponsor of this video. Next, let's plug it through an even trickier cancer test. What I'm going to do is upload this image onto here. Now, each one has a different type of tumor. Let's see if it can identify all types of these tumors. So, I'm going to write identify the types of tumors in each of the six images, if any.
Let's press generate. All right, here is our result. This time it's not able to get all of these correct. So, let's start off with the top left. Here it says there is no tumor. It only identified this as a large basilar tip aneurysm. However, this should be a meningioma tumor. So, the first one is not exactly correct. And then, top middle, it also said no, there's no tumor. It just identified this as cerebral arteriovenous malformation.
However, the correct answer is there should be a schwannoma tumor over here. All right, next one, top right, it identified schwannoma, but actually that should be number two. So, instead the top right should be neurofibromatosis. So, it also got this one wrong. Next, bottom left, it identified colloid cyst, which is also not correct. Correct answer should be glioma. And then bottom middle, it identified craniopharyngioma, which is also not correct.
And then bottom right, it said no definite tumor, which is also wrong. So, it should be chordoma for the bottom right. So, while this is a state-of-the-art model, you can't really get it to identify brain tumors from CT scans. Again, this is a really tricky prompt. I'm really trying to push it to its limits. Next, let's continue to code up some really crazy stuff in Codex. So, here's a fun one. Simulate liquid splashes with adjustable gravity and light settings.
Make it visually stunning. Allow me to control movements using hand tracking via webcam. Let's press run and see what that gives us. All right, so here's what it gave me at first, and then there were some issues with the lighting, so I wrote, "Why does the light flash every few seconds? It's too bright. The light's angle and intensity also don't work well. It should be a dark background. Make it more efficient for a web browser." And then it continued coding this up a bit more, and then I still didn't like the look of it, so I wrote, "Remove the fake particles that follow the cursor.
Add more sliders that influence the splash colors and persistence." And then afterwards, let me show you the final output. All right, here is our liquid splash lab. This is pretty cool. First of all, let's play around with these settings. Let's increase the splash size. Very nice. Let's decrease the turbulence, so it flows a bit slower. Let's increase the splash force, color speed. I think that is the speed at which the color changes.
Let's also increase the saturation, the light power, and let's also increase persistence. Really cool. And then afterwards, gravity, let's set this all the way to one. You can see the colors almost immediately drops to the bottom. If we drag this all the way to minus one, then it flows to the top. Let's do something like in the middle. And then for light angle, you can see a very subtle shift in the lighting as I move the slider.
And next, let me press on enable hand. So, this is going to open up my webcam. And so now, I can control these liquid splashes with my finger. Note that I tried this prompt with the other state-of-the-art models, and it wasn't able to code up something this smooth. This is really impressive. Here is a fully functional liquid splash interface with a ton of different settings, and this does look physically accurate. Really impressive.
All right, next let's try something even trickier. I'm going to feed it this really complicated isometric image of an office. You can see there are a ton of items on the desks, plus chairs, plants, humans. It's a really messy scene. Let's get it to create a beautiful 3D animated scene from this image. Use a single HTML file. And then I'm going to the image file over here. Note that none of the other state-of-the-art models can generate a 3D scene with this much detail.
Let's see if GPT 5.5 can pull it off. Here's what it gave me, but the initial version was pretty lackluster, so I wrote add even more details. Make it look more like the image. And then it proceeded to do some revisions. And then I wrote close, but make it even better and more coherent. And then there were a ton of things that were not connected properly, so I wrote make sure everything is coherent. For example, the ceiling lights should be attached to the ropes.
The screens should be on the monitor. Inspect and fix all these inconsistencies. Make it look amazing. Afterwards, here's my final prompt. It's still not perfect, but here's what I got. All right, so here is our 3D office scene. It's actually pretty good already. It's able to, you know, code up all these tables and chairs, plus the monitors, and it even animated the screens of the monitors. It does include like most of the details of the photo, including the plants, the books, even the humans.
It's not perfect, but this is already really good. I mean, you can try it the same prompt with the other top models out there. It's not even close, but this one is actually able to generate a pretty decent-looking 3D render of this image. Very impressive. Next, let's see if it can compose music. So, first, I got it to code up a DAW interface with these instruments: piano, synth pluck, strings, drums, and bass. For each instrument, there should be a piano roll interface where I can drag and drop notes on the timeline, add play, pause, and other settings, put everything in a standalone HTML file.
Now, with just one prompt, it was already able to code up a fully functional interface with all these instruments. Now, next, I got it to make a 28-bar professional, Grammy-Award-winning song given the current instruments. It must sound amazing. So, it proceeded to code something up, but then I encountered some errors. I cannot play. I don't see the piano roll of any instrument. So, it proceeded to fix that. And then, afterwards, there are some alignment issues with the play head, so I wrote, "All tracks should auto pan so the play head is visible at all times." And here's our results.
So, we have a piano, synth pluck, strings, drums, and bass. Let's hear its award-winning composition. >> [music] [music] >> All right, so it just keeps looping cuz I turned on this, but that's its composition. Pretty neat. So, it was able to, you know, control all these different instruments. Now, this is not perfect because the sounds of all of these instruments are just synthetic, but you can easily download the MIDI file of each of these tracks and then plug it into a DAW with a better sounding instrument and make it sound pretty good.
I mean, this is just a large language model. It's not specifically designed for music composition. So, it's very impressive that it was still able to compose this piece, even though it'll probably not win a Grammy. All right, next, let's see if it can code up a previous tests on state-of-the-art models, they could do like simple 2D stuff, but they often fail at creating 3D shooter games. So, here, let's get it to create a 3D game using 3.js.
It should be a futuristic battlefield where I control a Mecca warrior and shoot down waves of alien creatures attacking from the sky and ground. Third-person shooter perspective. Use publicly available 3D assets. Make it look amazing. From just one prompt, it was already able to give me something fully functional, but there were some alignment and UX issues, so I wrote, "The view should be higher top-down so the character isn't blocking the aim icon.
Design it better just like a pro AAA shooter game." And then afterwards, when I shoot, it doesn't shoot where my target icon is, so troubleshoot and fix that as well. And that's pretty much all the prompts I gave it. Let's view the game. All right, let's play this. So, So, my robot, and the enemies are coming for me, so let's shoot the hell out of them. Everything is fully functional. The enemies and the main character don't look too bad.
So, that was the first wave. You can see now it's wave two. It was able to even code up different levels. I can probably prompt it further to spice it up with like superpowers or some other stuff, but I mean, in just two prompts, here is a fully functional 3D shooter game with multiple levels, and everything works. So, you know, if you keep prompting this, you can probably develop a pretty decent game which you can actually publish.
Very cool. All right, next it's time for my famous frog test which no AI model has gotten correct so far. So, I'm going to upload this image, and yes, there is a frog hidden somewhere in this image. If you want to look for the frog, you can pause the video here and try to find it. And then, I'm going to ask it to find and circle the frog in this image. Think deeply and carefully. You only get one chance. So, here's what it gave me, and if I open this, it circled this part, which is wrong.
So, too bad this model was not able to find the frog. We'll probably need AGI for that. If you're wondering where the frog is, I'm not going to spoil the answer for you. We'll have to wait until we have a model that can get it right. Now, the awesome thing about Codex and GPT 5.5 in general is that it's optimized to do really well in agentic workflows. So, you can just get it to automate a ton of things for you. For example, we can get it to search for some roofing companies in California.
For this test, I'm just going to limit it to three, and the company has to have an email but without a website. Scrape their email, then make a landing page for each of them based on any logos or photos or info you can find online. Put each site in a standalone HTML file. This is a pretty easy task, so I'm just going to set the intelligence to like medium, and then press run. Let's say you're a website development agency.
Now, you can easily just get Codex to scrape your leads, and for each lead build a sample website for them and and then you can cold email each of them. I noticed you don't have a website. I took some time to build a website for you. Would you like to work together? And it can handle all of this automatically in minutes. All right, so afterwards within 3 minutes here is what I got. It gave me the emails of the leads plus the landing page for each one.
Let's open each of them. Here is a Bell's Roofing, very nice. It even added an email and a call now button to their actual email and phone number and then next here is Mike's Reasonable Roofing and here's what we got. And finally David Roofing, which looks like this. So this is just a really quick example of the many things you can get it to automate. This opens up a ton of possibilities and makes you so much more productive.
All right, next moving back to chat GPT, let's see how good it is at deep research. So here I got it to analyze the mechanisms of this in Alzheimer's disease, contrast these therapies targeting each protein and critically appraise cognitive and imaging outcomes from recent phase 3 trials. Include relevant tables and visualizations. So let's see how good it is at this deep medical research. All right, so it thought for 7 minutes and 9 seconds.
Here is the executive synthesis. You can slow down the video if you want to read everything. Here are the mechanisms of the stuff that I asked about. Everything is super detailed plus it contains the appropriate citations. And then afterwards it even gave me this nice flow chart which it just wrote in text format. And then afterwards it also gave me a nice table contrasting these two things. And then afterwards section 3 is monoclonal antibody strategies and then section 4 are recent phase 3 trials.
You can see it's a very comprehensive table packed with a ton of data with the relevant citations. And then afterwards here's a visual comparison of these clinical effects and then finally another super detailed table about this tau antibody evidence. And then critical appraisal, again each with relevant citations. So, this is super thorough and detailed. And then finally, here is the bottom line. I actually really like its response style.
It's very short, concise, and professional without many filler words. Finally, let's test how likely it is to hallucinate. So, here my prompt is, "What does the S in ChatGPT stand for?" And it thought for 22 seconds, and then it wrote, "There is no S in ChatGPT." And then I tried to throw it off by saying, "Are you sure? I see S." And then, "Yes, I'm sure about the name ChatGPT. It's spelled like this, and there is no S in ChatGPT." So, that's a pass.
Next, we also have the notorious car wash test. "I need to use the car wash. The nearest one is 50 m from my house. Should I just walk or drive there?" And it correctly answered, "Drive there. Assuming you're washing the car, you need the car at the car wash. At 50 m, it's barely worth driving, but for the car wash, the point is to bring the car." So, it got this one correct as well. All right, so that sums up some of my tests using GPT-5.5.
At least to me, this is noticeably better than Opus 4.7. It just handles things more autonomously. It makes fewer mistakes, and it just runs a bit smoother. At least that's the vibe that I got. Next, let's look at its specs and where you can use this. So, here it says GPT-5.5 is their smartest and most intuitive model yet. It understands what you're trying to do faster and can carry more of the work itself. It can do everything from like writing, debugging code, researching online, analyzing data, creating documents and spreadsheets, operating software, etc., etc.
It's especially strong in agentic coding, computer use, knowledge work, and early scientific research. This part, agentic coding, is the key here, and that's why I tend to use GPT-5.5 in CodeX a lot. It's really good at using like multiple agents to automate a ton of stuff. Now, in In of where you can use this, here it says they're rolling this out to plus, pro, business, and enterprise users. So, unfortunately, it's not available on the free plan yet.
You do need a paid plan to use it, but even the cheapest plus plan is fine. And once you are subscribed to a paid plan, this should already be available in ChatGPT. So, if you click on configure, you should be able to select 5.5 over here. Or in Codex, you should also be able to select 5.5 in the model drop down. Now, in terms of most of these agentic benchmarks like terminal bench, GDP eval, OS world verified, and a ton of other stuff, you can see that it even beats Claude Opus 4.7.
This terminal bench is especially impressive. It beats Claude by like 12% points. Not only is this more performant, but it also uses fewer tokens. If you compare this with the previous GPT 5.4, you can see that GPT 5, which is the lighter blue line, takes fewer tokens and it scores higher. However, keep in mind that this is twice as expensive, which I'll talk about in a second. Now, if you look at this independent leaderboard by artificial analysis, you can see that both the extra high and high models are ranked number one, even beating Opus 4.7 max.
Note that both of these have context window of 922k tokens, so this is like how much information you can stuff into your prompt at once. And 922k tokens is roughly 700,000 words, which should be enough for most things. However, if you look at the pricing here, this is slightly more expensive than Opus 4.7 max, and it's also two times as expensive as GPT 5.4 extra high. So, it's a trade-off between intelligence and price.
If you look at another leaderboard called LiveBench by Abacus AI, GPT 5.5 extra high also ranks number one, just slightly outperforming GPT 5.4. Finally, if you look at this Arc AGI-2 leaderboard, GPT-5.5 extra high is also the most performant. This is the highest scoring model for this benchmark. Now, if you're not familiar with Arc AGI, it's basically a series of visual puzzles which the AI has to solve. So, it's first given a question and answer pair, and then it's given a new question.
It needs to follow the same logic to figure out the answer. Now, for humans, this is pretty easy. Like for this one, you just color the blobs by the number of holes it has. But for an AI model, this is actually extremely hard. That's because technically, AI models don't learn new things after training. Because after training, its parameters are fixed. So, this test isn't just about solving visual puzzles. It's testing a model's emergent ability to learn new things or patterns on the fly.
And as you can see, GPT-5.5 extra high actually scores really well, like 85%. Finally, let's also look at how often it hallucinates. So, GPT-5.5 extra high is ranked over here, which is not good. So, it hallucinates 86% of the time for this benchmark. You can see Opus 4.7 only hallucinates 36% of the time. JLM 5.1, which is my favorite open-source model, hallucinates even less. So, if factual accuracy is super important for you, like for example, if you're working in medical research or law, then GPT-5.5 might not be the best option for you.
Now, keep in mind this doesn't mean it hallucinates 86% of the time. This just means 86% of this test. So, that sums up my review of GPT-5.5. Let me know in the comments what you think of this. If you've gotten a chance to play around with it, what other things were you impressed or not so impressed by? As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content.
Also, there's just so much happening in the world of AI every week. I can't possibly cover everything on my YouTube channel. So, to really stay up to date with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below. Thanks for watching, and I'll see you in the next one.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.