Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Search · @theAIsearch
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Search's most watched videos.
Most replayed moment at 10:28
3.1x that video's typical replay level
and then fix any errors that it sees. And so afterwards, it rendered the video successfully and that's pretty much it. In just one prompt, I didn't even need to prompt it further. Here's our final result. >> Four companies, one quarter, and a half-trillion-dollar bet on artificial
Said at 10:22
Most replayed moment at 15:21
4.1x that video's typical replay level
and deep agent for only $10 a month. This is way cheaper than if you paid for each tool separately. Definitely check out chat.llm that comes with deep agent in the description below. You can think of the residual connections not as a simple pipe carrying the signal forward,
Said at 15:15
Most replayed moment at 12:41
6.0x that video's typical replay level
other content, Higgs Field is a game changer that will supercharge your production workflow. Try it today using the link in the description below. Now, if we dive deeper, here's how it works in technical terms. They used something called a Markov head. In probability theory, a Markov process assumes that
Said at 12:33
The graph counts replays. It does not show where viewers stopped watching.
Words
3,350
Runtime
17:04
Speaking pace
196wpm
Reading time
14min
196 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
LTX 2.5 just came out and this is the fastest open source model you can use right now. In this video, we're going to go over its specs and new features. Plus, of course, I'm going to show you how to install it on your computer so you can use it for free and unlimited times offline. Let's jump right in. Now, this is an improvement over their previous LTX 2.3 model. So, let's go over some of the biggest changes. First of all, they introduced something called diffusion fidelity rendering. How this works is instead of spending
98 words, the words spoken in the first 30 seconds at 196 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 243 |
| Average words per sentence | 13.8 |
| Longest sentence | 43 words |
| Questions asked | 1 |
| Sentences containing a number | 44 |
Most used terms
Filler phrases
23 in total: like 13 · basically 7 · actually 2 · you know 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
LTX 2.5 just came out and this is the fastest open source model you can use right now. In this video, we're going to go over its specs and new features. Plus, of course, I'm going to show you how to install it on your computer so you can use it for free and unlimited times offline. Let's jump right in. Now, this is an improvement over their previous LTX 2.3 model. So, let's go over some of the biggest changes. First of all, they introduced something called diffusion fidelity rendering.
How this works is instead of spending the same amount of compute everywhere, it actually automatically allocates compute according to the scene complexity. So, for example, if the scene is really difficult with a lot of details or a lot of action, then it'll allocate more compute to that spot. But, if the scene is a lot simpler with fewer movements, then it's going to allocate less compute. So, this automatically helps you optimize efficiency.
Another new feature is you can now generate multi-shot videos. In other words, one generation could have multiple cuts of the scene at different angles, but everything would still remain consistent including the characters, the objects, and the overall scene. They also claim that this has much cleaner motion and better prompt understanding compared to the previous LTX 2.3. The nice thing about it is you can generate up to 4K in resolution and up to 50 frames per second.
You can safely generate videos of up to 20 seconds and you could potentially even extend it further depending on how much VRAM you have. And this thing is insanely fast. At least for me on my computer, it's more than two times faster than MiniMax H3. Plus, the nice thing about this is it supports existing LTX 2 LORAs. And there are already a ton of LORAs available that were created by the community that can help you generate different styles or different effects, etc.
So, this can be very powerful and flexible. Now, of course, a ton of you are wondering how this compares to MiniMax H3. I did a ton of direct comparisons in my last video, so I'll link to this in the description below if you want to check it out. So, next, let's go over how to install and use LTX 2.5. Here it says the minimum VRAM is 16 GB, but with some optimization, which I'll show you later in the video, you could also potentially run this with 12 GB or less.
Now, for this tutorial, we are going to use ComfyUI, which is the most popular platform for running open-source image and video generators offline. So, I'm going to assume you already have ComfyUI installed. If you don't, definitely see this video first for a full installation tutorial. All right, the first thing you need to do is to update ComfyUI to the latest version in order to see the workflows. So, in your root ComfyUI folder, simply click into the update folder, and then click on update comfyui.bat.
And this will proceed to update Comfy to the latest version. And then afterwards, it says press any key to continue, so let's press any key to exit out of this. And then next, we can start up ComfyUI. So, let me start it up. All right, afterwards, in your ComfyUI, on the left sidebar, simply click on templates, and then search for LTX 2.5. And you should see this. And if for whatever reason you don't see this workflow in your templates, I'll also link to this page where if you scroll down a bit, here it contains the workflow, which you can manually download and drag and drop onto your interface.
Now, there are several pro workflows with this crown icon. These are paid. This connects to their cloud service. So, what we're going to use instead are the free and offline ones, which are over here. So, let's first go over text-to-video. After we open this up, it's going to show us several errors. Specifically, we are missing some models. So, let's proceed to download all the models first. So, I'm going to link to this page in the description below.
We need to download several of these files. So, first, let's click into diffusion models, and then here is where you can download either a dev model or a distilled model. Now, the dev model requires around 20 to 30 steps to run one generation, so it's a bit slower. I would not recommend using this to generate videos. This is more for training LoRAs. So, if you just want to generate videos, it's better to go with a distilled model, which can generate a video in just like four to six steps.
Now, within the distilled LoRAs, there are different compressed versions. The full BF16 one is 42 GB in size, which will likely not fit for most of you. We also have a INT8 ComfyUI version, which is 22 GB. This should probably fit comfortably with like 16 GB of VRAM with some optimization. Or, if you have the right GPU, you can also go for this FP4 one, which is only 19 GB. For me, I'm going to download this INT8 one, so let's press download.
And this goes into ComfyUI in models, and then diffusion models. Let's click save. All right, afterwards, let's go back to the root folder, and then we also need to click on this latent upscale model, and let's download this spatial upscaler, which is around a GB in size. So, let's click on download, and this also goes in ComfyUI in models, and then in latent upscale models. Let's click save. All right, afterwards, let's go back to the root folder.
And then next, we need to also download the text encoder. So, here they were also updated the Gemma 4 text encoder, which is supposed to be better than the previous version for LTX 2.3. Again, we are given two different versions. This one is more compressed, it's only 16 GB, so let's download this one. And this goes in ComfyUI in models, and then in text encoders. Let's click save. Finally, we also need to download the VAE for this.
So, let's click into this VAE folder, and then we need to download the audio VAE, so let's click download. So, this goes into ComfyUI in models, and then VAE. And then afterwards, we also need to download one of these video VAEs. Now, this ComfyUI one is slightly smaller, so I'm going to go with this one. Let's click download, and this also goes in models, and then VAE. Let's click save. All right, those are all the files that you need to download to get started.
So, back in our ComfyUI workflow, after you've downloaded the models, simply press R to refresh your model list. And then for the model drop-downs, simply select the one that you downloaded. So, for me, for this one, I'm going to select this LTX 2.5 model. For video VAE, I'm going to select this one. For audio VAE, I'm going to select the audio VAE. And then, for the text encoder, I'm going to select Gemma 4. And then, for the spatial upscaler, I'm going to select this.
And then, for the prompt enhance model, you can just select any existing thing you have because we are going to turn this off. The prompt enhancer is going to take up more compute, and I don't think it's necessary. So, afterwards, after loading the models, the red outline and the errors should disappear. So, next, we can proceed to generate the video. Here are some additional settings. So, for prompt enhance, this basically uses whatever model you have here to enhance your prompt further.
But, of course, it's going to take up more time and compute, so I just tend to turn this off. And then, here's where you would set the duration. Let's just keep it at 5 seconds, but this can do up to 20 seconds safely. You can probably even do longer if your hardware can handle it. And then, here is the frame rate. Here is where we would set the aspect ratio and then the resolution. And then, up here is where we would enter our prompts.
I'm just going to leave it at this prompt. Now, if you click this icon in the corner, it will expand the workflow. And let me just go over how this works in simple terms. Now, up here is where we can use the optional prompt enhancer to edit the prompt further, but since we switched it off, this part would be disabled. It's actually first going to generate low resolution in the first pass, and then it's going to plug it through our upscaler.
Remember, we loaded a spatial upscaler over here. And that upscaler is basically going to upscale the video into our desired resolution. And that's what makes LTX generation so fast. It's because it first generates a low resolution video in its first pass, which is very quick, and then it uses an upscaler to generate the full resolution in the second pass. Anyway, that is how the workflow works. Next, let's press run to generate the video.
All right, so you can see this was very quick. It only took like 20 seconds to generate the video. And here's our result. All right, so that was text-to-video. Next, let's go over image-to-video. So, if you click on templates and then search for LCX 2.5 again, you should see this image-to-video workflow. So, let's click on this. And when you start this, it's going to show some errors again because we need to select the appropriate models.
So, let me do that really quickly. It's the exact same models as what we downloaded before. And then afterwards, here is where we can upload an image. So, let me upload this image of a cat rock band. And then over here is where we would input our prompt. So, let me enter this. And again, here's where we would select the aspect ratio and the resolution. Now, for me, since this is a vertical video, let's set this to 9 by 16.
And then over here is where we would set the duration. Again, we are going to set prompt enhance to off. Here's where we would select the frame rate. And that's pretty much it. And then if you expand the workflow, again, it looks the same as the text-to-image workflow. So, here is where we have the optional prompt enhancer, which we have disabled. First, it runs the video through a single low-resolution pass. And then it uses the upscaler to basically upscale the video to our desired resolution.
And that's pretty much it. Let's press run. All right, so again, that was incredibly quick. That only took like 20 seconds. And here's the result. >> [music and singing] >> Very nice. So, that is image-to-video. Next, let's go over the final workflow. So, again, I'm going to click on templates and then at the top here, type in LCX 2.5. And then let's click on this first frame, last frame workflow. So, again, similar to the last image-to-video workflow, but here we are going to enter one image for the first frame and one image for the last frame.
First, for each model, let me select my downloaded model as before. So, I'm going to do that really quickly. And then afterwards, again, we are going to set prompt enhance to false. Let's set the duration to five and for the width and height, let's just leave it at this. And then let me upload an image for the first frame and I'm going to upload this image for the last frame. All right, so we have these two images. So let me write this prompt.
So it's going to be a fast zoom through the city and then ending with this frame. If we expand the workflow, it's the same as before. Here at the top is the optional prompt enhancer which we disabled. Now interestingly for here there is no first pass at low resolution and then plugging it through the upscaler. There's no upscaler present in this workflow. So first frame last frame might take a bit longer than the previous workflows.
Anyway, let's press run and see what we get. All right, so that took around 30 seconds, which is 10 seconds more than just an image to video workflow, but still very manageable. That's what I love about LTX is that this is like two to three times faster than MiniMax H3. Anyway, here is our results. Now the awesome thing about LTX 2.5 is that this is compatible with previous LTX 2 Laura's. If you're not familiar with the term Laura, this is basically like a fine-tuned model by the community which can help you generate a certain character or action or style or effect.
For example, we have a Laura for fantasy realism. We have this Laura for better motion. We have this Laura for a K-pop dance. We have this Laura for transforming into a creature, I guess, or we have this retro 90s anime style Laura. In fact, let me download this and show you how to load it into LTX 2.5. So I'm going to click on download here and this goes in ComfyUI in models and then Laura's. Let's click save. All right, afterwards let me open up a text to video workflow so I can show you how to generate a simple video with this retro anime Laura.
So for my prompt, I'm going to write anime style a girl walking through a regular street in Tokyo. Pretty simple prompt. And then, if we expand the workflow, we basically need to add the Laura after this diffusion model. So, let me double-click anywhere on the interface, and then search for Laura. There are various load Laura nodes you can choose. I'm just going to go with the simplest one, which is load Laura by comfy, and let's set this somewhere here.
And here is where I can select the retro anime Laura, which I just downloaded. And then, for the strength, this is how much influence you want the Laura to have on your generation. Let's set this to something like 90%. So, basically, we need to connect this Laura to this load diffusion model, and then whatever goes after it. So, this goes all the way to this model input over here, as well as this model input over here.
Now, for most Lauras, it also requires a trigger word. So, for example, this retro anime Laura requires one of these trigger words, depending on, you know, what type of retro anime you want to generate. So, let's try this Akira one. I'm going to copy this, and then back to our workflow. For my prompt, I'm also going to input this trigger word. And that's pretty much it. Let's press run. All right, here's our result. It does look like retro anime style.
So, that's how you can load a Laura onto your workflow. All right, now, like I said, the official page says that the minimum VRAM is 16 GB. Even if you use the int8 conv rot model, this is like 22 GB in size. So, what if you have even lower VRAM? Well, fortunately, the community has already created more compressed GGUF versions of LTX 2.5. I recommend this one by Abby Ray. So, I'll link to this page in the description below.
And if you click on files and versions, note that they've released GGUFs of various compressions. The smallest one, Q3 small, is only 12.6 GB in size. So, this would likely fit within 12 GB of VRAM. So, let me just show you an example downloading the smallest one. I'm going to click on download, and this goes in ComfyUI in models, and then in unit. Let's click save. Afterwards, back in our ComfyUI, you can run this GGUF with any of the workflows.
Let me just show you a simple text-to-video example. Now, we just need to change one thing in this workflow, which is if you expand the workflow over here, instead of load diffusion model, we need to replace this node with the unit loader. So, let me double-click anywhere on the interface, and then search for unit, and we need to use this one, unit loader GGUF. So, let's click on this and place it here. Now, a nice trick you can do to take the same connections and drag it onto this new node is to hold down shift and then click on the connections and then move it over here.
And then, I'm going to do the same for this one, and that's pretty much it. Now, let me get rid of this load diffusion model node. I can just press Ctrl B to bypass this, and then if I escape to go back to the parent workflow, here for this model field, I can click on the drop-down and select my newly downloaded unit GGUF. And that's pretty much it. So, let's click run. And here's our result. As you can see, even the smallest Q3 version isn't too bad.
So, that's how you can run LTX 2.5 with potentially even lower VRAM. So, we've covered all the basics already. We covered text-to-video, image-to-video, first frame, last frame, how to load LoRAs, how to use GGUFs. Here are some more advanced stuff. So, this LTX upscaler can also be used with other models. For example, we can first generate a video using Minimax, and then link it through the LTX 2.5 upscaler to make this even higher resolution.
Now, it's hit or miss at times. Sometimes the upscaler is not really good, especially if you need to upscale high action scenes, but for like slow moving shots, then it is pretty decent. It's a bit too technical for this video, but if you're interested, I will link to this workflow by this user in the description below. Another powerful feature is this LTX director node. This is basically a mini video editor inside ComfyUI, so you can combine different workflows together like text to video, image to video, first frame, last frame, or even custom audio, and you can generate multiple clips and stitch them together.
Now, this was previously designed for LTX 2.3, but this also works with LTX 2.5. Again, this is very technical and beyond the scope of this tutorial, but if you're interested, I will link to this GitHub repo by What Dreams Cost, which contain all the instructions on how to download and set up LTX director. Anyway, that sums up my tutorial on LTX 2.5. This is incredibly efficient and definitely the fastest video generator out there.
This can do long videos and even handle 4K resolution. There's already a ton of lower support for this, so you can generate a ton of things with this. Let me know in the comments what you think of this. If you run into any errors with the installation, welcome to copy and paste the exact error message that you see in the comments below and I'll try to help you troubleshoot as much as possible. As always, I will be on the lookout for the top AI news and tools to share with you.
So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week. I can't possibly cover everything on my YouTube channel. So, to really stay up-to-date with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below. Thanks for watching and I'll see you in the next one.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.