Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Search · @theAIsearch
Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
10:225.1x the video's typical replay level
create any original film inside Higgs Field. It can be any story or genre and submit it for a chance to win cash prizes and other rewards. Try See Dance 2.5 in Higgs Field today using the link in the description below. So, for this tutorial, we are going to use a platform called ComfyUI. If you're
Said at 10:14
Most replayed moment #2
17:044.9x the video's typical replay level
breathe. Pictures start to move like the falling leaves. Video [music] drifting by at 24. Little animations on my bedroom. >> Now, to be fair, this ain't Suno quality. It's not as good as the most recent Suno models, but Minimax Music 3
Said at 16:57
Most replayed moment #3
15:594.4x the video's typical replay level
algorithms you can use to generate the song. For me, I'm just going to leave it at the default, and that's pretty much it. Let's click run. All right, here's my generation, and note that this took around 3 to 4 minutes on 16 GB of VRAM. Let's play this.
Said at 15:52
The graph counts replays. It does not show where viewers stopped watching.
Words
3,383
Runtime
21:35
Speaking pace
157wpm
Reading time
14min
157 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
The best open-source AI music generator is here. It's called Minimax music and not only can this create super clean and great sounding songs, but it's also incredibly tiny. The smallest model is only 2.5 GB in size, so this can fit on both consumer hardware. In this video, we're going to go over some demos and how to install it so you can run it for free and unlimited times offline. Let's jump right in. First, let me
79 words, the words spoken in the first 30 seconds at 157 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 249 |
| Average words per sentence | 13.6 |
| Longest sentence | 46 words |
| Questions asked | 8 |
| Sentences containing a number | 30 |
Most used terms
Filler phrases
32 in total: like 24 · basically 5 · kind of 1 · literally 1 · you know 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
The best open-source AI music generator is here. It's called Minimax music and not only can this create super clean and great sounding songs, but it's also incredibly tiny. The smallest model is only 2.5 GB in size, so this can fit on both consumer hardware. In this video, we're going to go over some demos and how to install it so you can run it for free and unlimited times offline. Let's jump right in. First, let me show you some demos so you can get a good sense of the range of music it can generate.
First, here's a pop rock example. The prompt and the song description is on the left and the lyrics are on the right. >> We were the kind who never stayed in one place, running through the city with the wind in our face. Trading every secret like a currency [music] of trust. Didn't know the moment when it started turning dust. I still [music] scroll back just to see your name, but the messages feel like another game.
Forever friends, we promised it one day [music] and we haven't seen each other since the following day. Guess forever's quicker than we thought it'd be. Faded like a sticker on [music] a teenage diary. Yeah, we said we'd never change. >> [music] >> But life got in the way. Now you're a highlight in a story I outgrew. A blurry little moment in a world that was new. Funny how the rhythm doesn't hit the same when the people that you dance [music] with walk off down another street.
I could call you up, but what would I say? Hey, remember us? It feels too far away. Forever friends, we [music] promised it one day and we haven't seen each other since the following day. Guess forever's quicker than [music] we thought it'd be. Faded like a sticker on a teenage diary. Yeah, we said we'd never change. >> [singing] >> But life [music] got in the way. >> Instead of pop, here's a more punk rock example. >> Don't waste your time on me.
You're already [singing] the only one who keeps me rock steady. And the voices in my head all assure me that that's what you said. You missed me and it fell short this time then your fading smile [music] keeps me whole for a while. The feeling of your hand in mine is something that I never want to forget about that summer. >> [music] >> I was in the ninth grade when I fell in love for the first time and her name was >> And here's a really smooth R&B neo-soul song. >> [music] [music] >> You see the best in me when I don't.
[music] You speak life when I lose my hope. You build me up, make me believe. [music] But you're not mine and that cuts deep. We talk in whispers, hide [music] our truth. Chasing something we [singing] can't prove. You say [music] I make you feel alive. But baby, we're [singing] living a lie. I'm torn between what's right and real. This love is something I can't conceal. You're [music and singing] my calm, my storm, my air.
The sweetest air I ever take. I love you, but I can't stay." >> It can also kind of do EDM, so here's an EDM progressive house generation. >> [music] >> I grab my pen, heart starts to race. >> [music] >> Time to escape, to claim my space. Each stroke of ink breathes life anew. Worlds awaken where dreams come true. [music] I'm weaving tales untold, characters bold, adventures unfold in colors of gold. Through ink and paper, [music] I carve my way.
In this realm of wonder, forever I stay. >> [music] >> From the shadows of mind, stories ignite. Panels bloom under [music] late night light. Lines and curves >> Next, let's see if it can do jazz. It can also do different languages, so let's try a Chinese example. >> [music] [singing] [music] [singing] [music] [singing] [music] [singing] >> Running to you baby >> [singing] [singing] [music] [singing] >> Now some of you folks are probably wondering if it can do metal, so here's a metal example. >> [music] [music] >> The moon's a blade [music] that cuts the sky.
Shadows crawling, [music] they creep, they lie. I hear the whispers, [music] sharp as glass. They're coming fast. They're coming fast. >> [music] >> It's a night battle, hearts >> [music] >> collide. Run or fight, there's no place to hide. Echoes screaming [music] across the night. It's a night battle under the >> [music] >> moonlight. Streetlights [music] flicker, a broken code. >> [music] >> Every step feels like it's slowed. >> Let me know if that sounds metal enough to you.
Finally, here's a fun country example. >> [music] >> We know those times when you feel like there's a sign there on your back. It says, "I don't mind if you kick me. Seems like everybody has." Things go from bad to worse. You think they can't get worse than that. [music] And then they do. >> [music] >> You step off the straight and narrow and you don't know where you are. You use the needle of your [music and singing] compass to sew up your broken heart. >> [music] >> Ask directions from a genie in a bottle of Jim Beam and she lies to you.
That's when you learn the truth. Met a man down [music] on the corner. Said he made >> Notice that it's even able to make the guy sing in a Southern country accent. Pretty impressive. All right, so those are some demos. Hopefully that gives you a good sense of what it can do. Next, let's go over how to install this. If you want to supercharge your content creation, definitely check out Higgs Field, the sponsor of this video.
They've just added the most capable video generator out there, Seed Dance 2.5. The biggest improvement from this model is that you can now generate up to 30 seconds of video in a single pass with multiple shots and an actual narrative with audio built in. You can also extend an existing generation with new shots while keeping the same characters, locations, pacing, and overall look consistent. What really stands out is the reference system.
You can feed it up to 50 references at once, including 30 images, 10 videos, and 10 audio files. So, you can provide your characters, environment, visual style, motion, and soundtrack all in one generation. You also get much more control over editing. For example, you can specify exactly what happens during different timestamps, change just one section without affecting the rest of the video or move the same performance into a completely different environment or even change the camera angle while preserving the characters and action.
See Dance 2.5 supports text to video, image to video, video to video, and other references giving you ultimate flexibility on your video creation. And right now is the best time to try See Dance 2.5 on Higgs Field because they are offering unlimited See Dance 2.5 for up to 33 days. Terms and conditions apply. And if you're up for the challenge, check out the Higgs Field Global Film Festival, which has a massive $1 million prize pool.
Simply create any original film inside Higgs Field. It can be any story or genre and submit it for a chance to win cash prizes and other rewards. Try See Dance 2.5 in Higgs Field today using the link in the description below. So, for this tutorial, we are going to use a platform called ComfyUI. If you're not familiar with it, this is basically the most popular platform for running open-source image, video, and audio generators offline on your computer.
It's completely free for you to use. In fact, if you're not familiar with ComfyUI, definitely see this video first where I go over how to install and use it. All right, assuming you do have ComfyUI, the first thing you should do is update to the latest version. So, in your ComfyUI folder, simply click into the update folder and then click on update_comfyui.bat and it should proceed to update to the latest version. Afterwards, it says press any key to continue, so let's press any key to exit out of the terminal and then next we can proceed to start up ComfyUI.
After you've opened up ComfyUI, simply click on templates in the left sidebar and then at the top here search for music and you should see this Minimax text to music. So, let's click on this and here's what the workflow looks like. Now, you first need to download several files for this to work. And by the way, if you don't see this workflow, I will also link to this page where you can manually download the workflow and then just drag and drop it onto your interface.
So, first over here we need to download a few models to get this to work. The first model is the unit, so I will link to this page in the description below. If you click on diffusion models, here is where you can download one of these models. The full model is only 9.8 GB in size, so even this should fit on like most consumer GPUs, which is great. There's also an FP6 version, which is 4.9 GB in size, and then an even smaller int8 Convo version, which is only 2.5 GB in size.
For me, I'm going to download this FP16 one, so let's click download and this goes in ComfyUI in models and then in diffusion models. Let's click save. And then afterwards, we also need to download this text encoder. So, in this text encoders folder, again we have three for you to choose from. The smallest int8 one is only 9.2 GB in size. So, that's what I'm going to download. Let me click on download and this goes in ComfyUI in models and then in text encoders.
Let's click save. All right, finally we also need to download the VAE for this. So, let's click into this folder and then there's only one VAE, which is 217 MB in size. Let's click on download and this goes in ComfyUI in models and then VAE. And that's pretty much it. After you've downloaded all these models, simply press R to refresh your model list and then select the model that you just downloaded from the drop down.
So, let me do that real quick. For the text encoder, I'm going to select this one. For VAE, I'm going to select this one. All right, so that's all you need to do. Now, next we can start generating this song. First of all, note that for now we only have a text-to-music workflow. There's no music-to-music. You can't create covers yet. I'll talk more about how you can do that later in the video. Now, the first input field is where you would input a description of the song.
Here it's recommended that you structure this field into three sections: a global metadata, which contains like the genre, the overall feel, plus like the pace and the key. So, for example, here we're going to input a lo-fi hip hop chill hop with jazzy extensions. It's going to be laid-back and dreamy throughout. A gentle warm drift with subtle night glow that deepens in the middle, etc., etc. And then for the second section, we should enter the vocal details.
How should the voice sound? Should it be male or female? Should it be high-pitched, low-pitched? Should it sound smooth or coarse? What should the timbre sound like? And then afterwards, the third section should be the arrangement. So, you can enter details about, you know, the exact instruments and what types of instruments to use. You can also enter details about the intro, the verses, the bridge and outro, etc. Now again, you don't have to come up with all of this from scratch.
If you're not sure, you can just copy this prompt and add it to ChatGPT and get it to write it out for you. Now, the next field is where we can enter the lyrics. And as with the other Frontier music generators, you can add meta tags like intro, verse, bridge, chorus, outro, etc. You can also add like secondary vocals in brackets like this. It even works with mhm's and oh's and other sounds as well as pauses like this.
And then after you enter in your lyrics, here is where you set the maximum duration in seconds. So, this supports generations up to 5 minutes long. So, you could potentially set this all the way to 300 seconds. For me, I'm just going to leave it at 1 minute. And then the seed is basically the unique ID of every song. So, if you leave all the settings the same and you use the same seed, you're going to get the exact same generation as before.
If you use a different seed, then you're going to get a different generation. And then finally down here for tiled and code, it explains what this is over here. So, basically, if you enable this, it's going to cut VRAM usage. So, this is helpful if you're generating very long songs on or 24 GB, then you can just turn this off to make it faster and better quality. So, since I do have enough, I'm going to turn this off.
And that's about it. Now, you could expand this workflow further by clicking this corner over here. And here you can also set the batch size, so like how many songs to generate at once. Let's just leave it at one. And then here's the classic K sampler, so you can play around with the step count. In general, the more steps you have, the higher quality the song will be, but it'll take longer to generate and vice versa.
And then CFG is like how literally you want the AI to follow your prompt. So, if you find that your generation isn't really following your description, you can try bumping up the CFG to see if it's more faithful. And then here are the different algorithms you can use to generate the song. For me, I'm just going to leave it at the default, and that's pretty much it. Let's click run. All right, here's my generation, and note that this took around 3 to 4 minutes on 16 GB of VRAM.
Let's play this. >> [music] >> Midnight and the canvas glows. Dragging little lines where the current flows. [music] Type of quiet dream, let the sampler drift. Noising to a picture like the fog just lifts. 20 slow steps, I'm in no hurry now. Latents turning colors and I don't know how. Every render is like a Polaroid I found. Soft focused memories, no sound. >> [music] [music] >> Cue another frame, let the motion breathe.
Pictures start to move like the falling leaves. Video [music] drifting by at 24. Little animations on my bedroom. >> Now, to be fair, this ain't Suno quality. It's not as good as the most recent Suno models, but Minimax Music 3 arrives at a very important time. That's because a few days ago Suno announced additional restrictions for their users. In particular, even paid pro users are limited to only 20 downloads per month starting September 3rd.
Premier users can get 60 downloads per month, and you can only use your songs commercially if they were generated with a paid plan. Now, 20 downloads per month isn't bad, but it has received quite some backlash among Suno users recently. Suno is also known to add watermarks or tracking to their generations. So, if you distribute your Suno generations off platform, they could potentially identify and track these generations.
All right. Now, like I said, currently this Mini Max Music 3 only supports text to music. But, what if you want to edit or inpaint an existing track, or maybe take reference from a song and do a cover on it? Well, fortunately, there are open-source tools for those features already. So, I've covered another tool called A Step 1.5 XL, and this allows you to inpaint existing songs or make music in the style of an existing song.
I already did a full tutorial on this, so see this video if you want to learn more. Another really cool open-source tool for music generation is called Foundation 1. And this basically lets you generate individual loops. You can take a loop as reference and generates additional tracks on top of that. And so, you can create multiple tracks and stack them together to create a full song. And of course, you can also import these tracks into a DAW to edit them further.
You can also convert each track into MIDI notes, so you can even adjust the individual notes afterwards yourself. And speaking of MIDI, we have another open-source tool called Muse Scriber, which basically allows you to upload a song, and it'll dissect the song into separate instruments and figure out the MIDI notes for each track, including the vocals. In fact, they released a free Hugging Face Space, so let's open this up and try an example.
Let's drop a song into this, and let me play you the song first. >> Every tear [music and singing] I cried made me realize I'm breaking free tonight from [music] your [singing] lies. >> And then down here we also have some event settings, so you can manually choose the instruments that are present in the song to improve the accuracy, but I'm just going to leave this all blank and get it to automatically select the instruments itself.
And that's pretty much it. Let's click on transcribe. You can see this is fairly quick. It only takes a few seconds for this to complete. And afterwards, it has detected a voice track and an acoustic piano track. And here's what the MIDI sounds like. >> [music] >> And there you go. It's as simple as that. There are some errors with the prediction during some parts of the track, especially at the end there, but you can import this into a DAW and adjust the notes further and of course remix this however you want.
So this is a really useful tool if you want to take any song and take elements of it and reverse engineer it. If you click on this hugging face link, note that they've released three different variants of this, a large, medium, and small variant. The large one is fairly tiny at only 5.5 GB in size, so even this one should fit on most consumer devices. And then the small one is only 0.1 billion parameters and this one is only 412 MB in size, so this could even potentially fit on just your phone.
So that's Muse Scriptor. This is really useful if you want to dissect a full song into the notes of each individual track. Anyway, that sums up my review of MiniMax Music 3 and some other open source tools that you can use for music generation. I hope you found it helpful and if you run into any errors during the installation, welcome to copy and paste the exact error message that you see in the comments below and I'll try to help you troubleshoot as much as possible.
As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week. I can't possibly cover everything on my YouTube channel. So, to really stay up-to-date with all that's going on in AI, be sure to subscribe to my free weekly newsletter.
The link to that will be in the description below. Thanks for watching and I'll see you in the next one.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.