Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Alex Ziskind · @AZisk
Words
2,693
Runtime
14:11
Speaking pace
190wpm
Reading time
11min
190 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
This is a 2.8 bill, sorry, this is a 2.8 trillion parameter model and it looks like it's going to be the biggest one for a while. Kimmy K3, 1.56 terabytes on disk. It doesn't fit on a single Mac Studio, not even in Q4, not even in Q2. That's the quantization I'm talking about when you shrink the model from its original weights. So, I wired up four Mac Studios together, each with 512 GB for a total of 2 TB of unified memory. 2 TB minus 1.56 Yeah, we're good. And I wanted
95 words, the words spoken in the first 30 seconds at 190 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 282 |
| Average words per sentence | 9.5 |
| Longest sentence | 43 words |
| Questions asked | 19 |
| Sentences containing a number | 44 |
Most used terms
Filler phrases
36 in total: like 11 · actually 7 · kind of 6 · I mean 3 · basically 3 · you know 3 · right? 2 · uh 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
This is a 2.8 bill, sorry, this is a 2.8 trillion parameter model and it looks like it's going to be the biggest one for a while. Kimmy K3, 1.56 terabytes on disk. It doesn't fit on a single Mac Studio, not even in Q4, not even in Q2. That's the quantization I'm talking about when you shrink the model from its original weights. So, I wired up four Mac Studios together, each with 512 GB for a total of 2 TB of unified memory. 2 TB minus 1.56 Yeah, we're good.
And I wanted to see if it can beat the new Abacus AI supercomputer. I'm going to have Kimmy K3 build a web app, a front-end web app, and I'm going to have supercomputer do the same. It's local hardware versus cloud death match, round two. Let's go. So, last time I did this it was one Mac Studio, 512 gigs against Abacus deep agent. And a lot of you mentioned in the comments, "Alex, that's not a real local setup." Fair enough.
So, this time I brought four Mac top and yeah, it's running on it right now charted across all four Mac Studios connected with RDMA. Mac Studio 1, 208 GB used, almost 200 used on Mac Studio 2, 177 on Mac 3, and 215 on Mac 4. For context, Kimmy K2 was 1 trillion parameters and that already felt pretty ridiculous. Kimmy K3's almost three times that. And the more parameters you have, the more the model can actually do. It's smarter.
That's the whole trade-off. A 1 billion parameter model gives you gibberish. No that's Now, even though 1.56 TB in the original 8-bit precision, I'm still running a smaller quantization, but it's unpruned. I'm running all 896 experts. That's a total of 817 GB of weights sitting on disk. Now, you might think, "Oh, each one of those is 512, you can fit that on two, right?" Well, not exactly because I'm giving it a large amount of context, you're going to need a lot more space.
And you could say, "Well, why not use three then? Because tensor parallelism splits the model dimensions across nodes. And those dimensions have to be divided evenly. So, it's either one, two, or four, or eight. Kimik4, maybe next year. And all four of these are connected through Thunderbolt 5 mesh on the back here. Six cables total connecting every machine to every other machine. I made a whole video describing how I did that and how I ran it with MLX distributed.
You can check that out. I'll link to it down below. All right. Let's see your pretty face. Now, let's meet the other contender. This is Abacus AI, and they have a new tool called Supercomputer. You've probably already seen that they sponsored some of my content in the past, and they're sponsoring this video, too. But, I keep coming back to it on my own time because essentially it gives me one login to be able to do and use every frontier model out there on day one.
There's over 100 models here, and route LLM basically routes you to the right one based on your context. So, there's GPT 5.6 Soul. That's pretty new as of this video. Llama 2 5, Cloud Fable 5, Grok 4.5. You know, they got image generation and things like that. They also have Kimik3 here, too. DeepSeek V4 Flash and Pro GLM 5.2. Open-source models, close-source models, everything. Kind of cool cuz you're not locked into a single ecosystem.
Now, what's different about Supercomputer is that it's actually a machine, a virtual machine. It's not just a chat window. It's a cloud box that's always on that's also happens to be integrated into their agent. You don't wait for a cold start. You don't wait for a spin-up. You never see container provisioning because it's always on. Scheduled tasks, databases, storage, GitHub integration, SSH. And also, they got Hermes agent now.
So, here I've got Hermes running, and I can change models to whatever I want. So, I've selected Kimik3. I can select whatever other model I want to talk to directly, like GPT 5 Luna. Say hi to Luna. So, you got Hermes. You don't need to install it locally. It's always on and running over there. I'll talk about the pricing in a bit. So, the fight is pretty simple. They're going to get the same prompt, both sides, and I got the prompt as a gist.
You can check it out. I'll link to it down below. It's a pretty extensive long prompt, but it's specifically front end for a silicon compute exchange front end. I want to make sure we're using TypeScript with the latest and greatest front end tech. We're using mock data, not real data, but this can easily be swapped out later. Dark mode, microinteractions, accessibility, we should have a pretty decent-looking site in the end, hopefully.
It's going to have some unit testing, and it needs to check its own work. So, this is the agentic part of it. It's going to be running all these commands to be able to verify and validate that it's working. Okay, let's get it. Raw copy. I'm going to head over to supercomputer and just paste this in into the agent. Now, here it's going to be set to auto, but you know what? I'm going to go with Opus 5 high and GPT 5.6 soul.
Oh, yeah, you have access to the full virtual machine here on the right, too. You got the files that it's going to generate, access [music] directly to the terminal, and access to the desktop. Now, I don't want to give it a leg up, so I'm going to go and start this off in my Mac Studio cluster, too. And the way I set this up locally is I just have Open Code, which is the local agent, pretty popular, but you can point it to any local model that you want.
And I've pointed it to the 4x Mac Studio cluster running Kimiko 3. I actually made it a separate video on how to set this up for members of the channel. By the way, thank you to the members for supporting the channel. Let's paste in that prompt, and boom. There we go. [music] It's doing stuff, I hope. How do I know? Here's all four of them, and the GPU usage is pegging 100% on all four of these machines. Do I hear it?
No, I don't hear it. Do I feel it? It's warm. And if I put my ear right up into it, then I'll hear a little little twinkle twinkle little star kind of sound in there. Maybe the fans will kick up later, but it's working. Do we have anything built yet? No, it's thinking. Where is my output? All right. All right, it's doing it. I'm going to start the other one. Boom, I'll build that for you right now. This will basically do the same thing.
It's building it. Let's pop this open on the right so we can see what's happening. By the way, this is the first time I'm actually running that prompt. I don't know what's going to happen. Hopefully it works. It says if you connect your GitHub, I can work with your repositories directly. That's pretty cool. Cloning, pushing, commits, and pull requests. Want me to set that up? Go ahead. So now Abacus AI agent takes over and we've seen it at work before, but the agent is now well integrated into the supercomputer concept that's always on and it's going to be deploying to this VM.
Building a production-ready Next.js web app. Now compared to some of the other builds I I've done here on the channel with Nvidia and DGX Sparks, the prefill or the prompt processing stage of inference is something that Max are not as good at as, for example, Nvidia GPUs. Macs are really good at generating once they've processed because they have extremely good high memory bandwidth. And with Kimiko 3, I measured about a 238 tokens per second chewing through the prompt, which is not super fast.
Oh. Now I can hear it. There's work to be had. Generation happens at about 14.7 tokens per second with Kimiko 3 here. It's not going to be fast, folks, but it is a big model. What's happening on the Abacus side? Well, we already have some HTML generated. Now while it's doing that, we can take a little peek under the hood and see what's happening. So I'm going to go in here and let's do uh this is just Ubuntu over here.
We got skills. This is just a generic skills repo. It doesn't have any custom skills installed. So there is Silicon Exchange. Let's go in there. That's our app. Next.js space. And yeah, there's our scripts and Tailwind config, TypeScript configuration, Next.js config. I wonder what's in that style guide. I wonder if this is something that it generated just now based on my prompt or just a generic one for Next.js projects.
Anybody know? >> [snorts] >> Please, sir, I want a token. Now, nobody's going to believe me that this Kimmy actually works. I swear, I tried it. Didn't try this prompt, but I tried using it and it worked fine. Maybe this prompt is just too hard. Now, let's take that 14 tokens per second figure. We can use something called Abacus AI Desktop, which is their desktop companion app. It's got chat, co-work, and code. And this is basically like a local agent.
They have a CLI, too, but they have a graphical interface. It's kind of like Codex. The difference is you can pick the model that you want right here from this drop-down. Fable 5 is there, Opus 5, GPT 5.6, Soul, Terra, Luna, Grok is there, and Kimmy K3 is there, too. So, let's see what we get here. Of course, I'm using code here, so I'm going to need to select a workspace. Write me a simple JavaScript Node application for adding two numbers.
Boom. Kimmy K3. Now, this is not using my Kimmy K3. This is using the Kimmy K3 that Abacus uses. But it's writing a local application. So, we need to accept a few things here. Allow it. Let's do a bypass, and there it goes. Now, that is pretty fast. I think it's done. Here you can access the VS Code extension, browser extension. So, this desktop app has a lot of extra abilities that it adds. Let's see. Node, add five and four.
Nine. Is that right? Think so. What's happening over here with our Kimmy? Oh, no token yet. Well, it's been a while, so I had to check what's going on, and it wasn't printing out any tokens at all. So, I'm going back into it to see if I can even get anything out of it. I'm going to just say hi. Does look like it's working. The GPUs are at 100%, so yeah. Hi, how can I help you today? It worked. Maybe I just wasn't patient enough for such a large prompt.
Let's do this. Create a simple web app with a text box and a button. Well, look at that. It just printed out index.html for me. Only took about 3 minutes to 147 lines. It's a beautiful, beautiful app. Wow. Test, submit, and it works. Still printing out. It's not done yet. I've created a simple web app for you. It's It's doing the explainer. Thank you. Thank you for that, Kimmy. I mean, you did a fantastic job. I just wish you did it faster.
Okay, after all that, I couldn't just leave it like that. I couldn't leave you hanging. So, I stopped recording, had some coffee, went back and gave Kimmy K3 the full prompt again. But, this time I didn't sit there and watch it because, you know, watching something It's kind of like watching grass grow. You don't think anything's happening, but something is. I had to walk away and have more coffee. Guess what happened 4 hours later?
We have a full app. Here it is. And it looks pretty good. This was done by Kimmy on these machines. There's the browse. We have live filtering, reset all filters, inspect, compare. We have charts. We have scheduling. This is a beautiful interface. And it can do light mode. It can do dark mode. Sorry about that. What do you think? I think it did a fantastic job here. So, the question, can you do this locally? The answer is yes.
Yes, you can. Let's take a look at what Abacus came up with. It looks like it's actually finished. So, let's pop open our VM here. Still looking at exchange, and then what's the name of that? Next.js app. And we're going to do npm run dev. Okay, it built and started it at this URL. Let's take a look at that. Okay. Wow. Can't get too excited cuz it is a sponsored video. Wow. Oh my gosh. >> [laughter] >> I mean, it is pretty cool.
It's very cool. It's a beautiful application. The design is quite something. The dashboard is beautiful. You got filters here on the left doing live filtering with animation. This is from a single prompt. I mean, this is not hooked up to a back end, but it wouldn't be that difficult to take it and extend it there. Your reservations, compare boards, browse inventory goes back to that page. Wait, what does this do over here?
Let's do Apple M3 Ultra available. Simple web app. Look, I'm sure the capability is there of this model with the right hardware. I'm capable to run it. It's just not very fast. And I like to code. And the way I use the agents is to do these one-off tools that I use. My GitHub lately has been a bunch of different tools that I implemented using the help of agents. These are things that I want to knock out quickly and get working.
And for that kind of stuff, Abacus supercomputer is perfect. Same prompt, same app, both finished it. 15 minutes against 4 hours. Now, each one of those Mac Studios was about $16,000, but now I think they're not for sale anymore, so it's a lot more. And I don't know how much the new ones are going to cost, but you get the idea. It's expensive hardware. You buy it one time, then it's yours, and then you do whatever you want with it.
Abacus starts at 10 bucks a month. Actually, $7 for the first month now. You could run that subscription for over 300 years before you spend that kind of money. And you can use their web UI, you can use the desktop application that I showed you. You can use their APIs. You're also not stuck at one price because there's custom routers that let you push the easy work to cheaper models and save the expensive ones for the harder stuff.
So, the bottom line is, and it's not what sponsors usually want to hear, Yes, you can do all the stuff locally, and that's pretty remarkable. That wasn't possible just a year ago. But, if you want high quality, and you want it done fast, and you want to be able to host it all in one shot, Abacus's cloud solution wins this one. If your data can't leave the building, or you just love running this stuff yourself, like I do, Cluster.
There's definitely a market for both, not one or the other. To see me build and configure that cluster and all the details of how I run it, watch this video right here. And watch this video when I compare a Mac Studio against Abacus's AI. By the way, that website is still up and running. >> [music] >> Thanks for watching, and I'll see you next time. >> [music]
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.