Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Latent Space · @LatentSpacePod
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in Latent Space's most watched videos.
Most replayed moment at 24:22
5.3x that video's typical replay level
you know, um it feels like I now just have an army of, you know, really dumb I O I like formalists. Yeah. And Yacoub was like, that feels like already the situation I'm in. >> [laughter] >> You're like, so uh yeah, no, he's he's
Said at 24:15
The graph counts replays. It does not show where viewers stopped watching.
Words
7,970
Runtime
44:02
Speaking pace
181wpm
Reading time
33min
181 words per minute, the same as the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
What used to be considered fast at like 1,000 100 or or 200 tokens per second is quickly becoming the new batch mode. And so I think that you know, there's a place for that. It's great for prompt processing. It's great for very very parallel workloads. Um >> But my eval's take 20 hours. I can hit a button and run it in 2 hours. >> Exactly. Exactly, right? [laughter] But but quickly, you know, being able to do prompt processing isn't isn't quite enough, right? That's [music] a big
91 words, the words spoken in the first 30 seconds at 181 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 482 |
| Average words per sentence | 16.5 |
| Longest sentence | 397 words |
| Questions asked | 134 |
| Sentences containing a number | 43 |
Most used terms
Filler phrases
622 in total: you know 201 · right? 101 · like 99 · um 80 · uh 79 · kind of 22 · I mean 18 · actually 15 · basically 3 · literally 2 · sort of 2.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
What used to be considered fast at like 1,000 100 or or 200 tokens per second is quickly becoming the new batch mode. And so I think that you know, there's a place for that. It's great for prompt processing. It's great for very very parallel workloads. Um >> But my eval's take 20 hours. I can hit a button and run it in 2 hours. >> Exactly. Exactly, right? [laughter] But but quickly, you know, being able to do prompt processing isn't isn't quite enough, right?
That's [music] a big part of the market, but that's not quite enough. >> Before we get into today's episode, I just have this small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content. We've been approached by sponsors on an almost daily basis, but fortunately enough of you actually subscribe to us to keep all this sustainable without ads.
And we want to keep it that way. But I just have one favor to ask all of you. The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you, and it means absolutely everything to me and my team that works so hard to bring the In Space to you each and every week. If you do it, I promise you we'll never stop working to make the show even better. Now, let's get into it.
Okay, we're here at Cerebras HQ with CTO Sean Lee. Welcome. Uh >> Thank you. Thank you for having us. >> Yeah, no, thank you for having me. >> Uh and it is the day after Hot Chips. Lots of things launching. Even like Apple launched >> [laughter] >> their M6 stuff. Uh but you guys obviously had We were at Supernova. You talked about CS-4. We're going to talk a little bit about CS-5. Open AI talks about jalapeno. All the all the hot stuff in Hot Chips.
What's your take this year having been in this industry? >> I think I think Well, first of all, thank you for having me. Um and and and you're absolutely right, it's a super exciting time right now. Um, Hot Chips is a is is a is really good, you know, time when the whole community comes together. Um, and you know, what what used to be, you know, I still remember going to Hot Chips 10, 20 years ago when it's a bunch of like, you know, computer architect, you know, geeks, you know, geeking out on, you know, speeds and feeds.
And and now it's like this is where, you know, industry-changing hardware is being revealed. And so, um, you know, it's come a long ways. It's a really exciting, you know, community. And I think one of the main takeaways that that that I had after, you know, after leaving Hot Chips yesterday is like, what an exciting time it is for the the chip design and the hardware industry. It's uh, you know, this is really the golden age of uh, of hardware right now.
Um, and you know, like you said, I've been in this industry for some time. Um, this is a very very unique time. Not just because, you know, AI is taking off and you know, and there's uh, you know, obviously a lot of uh, people doing their own hardware, but the amount of innovation that's happening now across multiple fronts, we've never seen in the history of uh, of the semiconductor industry this level of innovation across all the levels, right?
In the chip, uh, the interconnect, the system design, the software, the optics, the you know, the methodology, the tooling, people are pushing across the board. And so, like, you know, as a technologist and and, you know, a a computer architecture geek at heart, it's like really really fun to see this community and this industry at this stage right now. And when we and we felt that like at at Hot Chips. And everybody I talked to, it's like that that buzz is there.
It's it's amazing. >> We want to go through the chips. Let's, you know, talk about your chip first. What's what's most exciting? CS-4 came out some crazy numbers here, 30x faster than GPUs, but uh, you know, walk us through it, new with the new CS-4? >> Sure. So, you know, we we we we designed uh our next generation CS-4 architecture um with a new brand new system platform with the goal really um to make wafer scale uh mainstream, to make it uh uh to bring it to hyper scale, and to solve uh a lot of the the the high density um data center problems that you know we're we're we're facing every single day.
So, we've designed this uh a modular platform uh that provides twice the amount of power to the wafer than we have in our previous generation, um twice the amount of uh interconnect bandwidth, half the latency, um and all of it is done at the system uh architectural level. And what you see here is the result, is that by being able to provide significantly more power to the to the wafers by you know, packing more density into the rack, we're able to drive up the performance even further.
Um you know, we like to say that Cerebras is uh with our current product already the undisputed leader in, you know, ultra fast inference. Um and and what's amazing is this CS-4 product will actually push that frontier even further by taking it another two times faster. Um and this is a a perfect example, right? We here in this in this demo that we gave uh at Hot Chips, um we're showing uh GPT-J uh running at over 4,000 uh 400 TPS, which is just mind-blowing.
It's like it's it it it almost feels like it's fake, right? Um and uh and and we think that this is going to really revolutionize the entire, you know, industry because, you know, not only are we able to, you know, continue to push the frontier on how fast these models can run, um it will enable all sorts of new applications. Um, you know, uh, user experiences start to become extremely different, right? What used to be batch and offline applications start to become real time.
Um, and, you know, all of this, you know, growth and new agentic flows and agentic, uh, frameworks now also mean that, you know, you're you're sitting there waiting for these agents to go over and over and over, uh, in their agent agentic loop. And all of a sudden, if you're running uh, your model at over 4,000 tokens per second, now the you can do, you know, more agentic, uh, loops, you can do more reasoning. Ultimately, you get significantly more capable, more intelligent agents.
And so, this is what really excites us about this. Not just the fact that you wish to show these really, really big numbers, which is also really cool, um, but but the fact that you can, you know, really start to do things that you can't really do, uh, otherwise and and and start to enable brand new capabilities. Um, that that's that's what makes me so excited about this this reveal that we had yesterday. >> Yeah, I think it's also very rewarding that you guys started on this journey, the big chip journey.
Uh, you know, we first scale everything that everything that you've always said, it just like people didn't take it as seriously until they had to. And then now they're they're really, really taking it very super seriously, right? >> Yeah. >> Um, I I I do feel like as, you know, co-founder Zetian, the the architect of this whole thing, how are you guys approaching the new generations post like AI boom? Like I imagine CS 123 was a bit more, uh, sort of, uh, calm. >> [laughter] >> Now like literally OpenAI is like launching soup that the the ultra fast mode with you guys.
And like it it is a matter of like I don't know, like you're you're you're co-designing your model with the chip almost. >> No, absolutely. I I feel like the main shift that has happened over the last few years, um, for us is that you know, the first generation chip primarily was a technology demonstration, right? We had to demonstrate to the world, to ourselves, that, you know, you can actually build such a thing. Um and uh you know, we we we told a lot of people that this is the future and most of them were just like, you guys are insane.
This is not possible. Um and so, you know, at the at the at the beginning it was really about, you know, figuring out the foundational technologies to make it possible. And what we're now seeing is now that we have in fact demonstrated that this can be built and it can have the kind of performance that, you know, that this type of architecture can have, what are all the things that it can enable, right? And you you know, you alluded to, you know, our recent ultra-fast launch with OpenAI.
That's a perfect example, right? Um we're running uh uh you know, frontier level, one of the most intelligent models, right? OpenAI's largest, most capable, most intelligent model, right now at 14 times faster than their normal, you know, GPU speeds. And it's becoming this completely transformational thing when you can combine the most intelligent models with the speed that the hardware brings. And this transition from technology demonstration to, you know, solving real problems, um enabling new new real capabilities, is this transition that we as a company are going through.
And how we're coping with that right now is we're really shifting our focus to ensuring that we can make this scale, we can bring up bring the capacity that we need, we can uh uh run the models that um you know, the users ultimately, you know, care about. Um And then we continue this flywheel by making sure that we continue to put out faster hardware and keep pushing the frontier and that's that's kind of you know how things have shifted for us and it's I would say like you know unfolding in a way that is you know almost better than we could have imagined right is this perfect confluence of like you know having the hardware with the capabilities demonstrated now matching with the capability of the models and the applications and and so our goal right now is just you know scale this thing as fast as we can. >> And so maybe we'll talk about yesterday. >> Yeah, sure. >> Previewing CS5.
Can you recap for people who are maybe not not yet caught up and you know maybe it's our new to the story for the first time as well. >> Yeah, no no problem. So I mentioned that our CS4 system that we just launched is based off of a brand new rack scale platform that we call the Nexus platform. And as I mentioned the what's special about this platform is the modularity. It has a modular power supply in the front. It has a modular server in the back that we call the backpack.
And this is a platform that enables that 2x performance that we just all saw the numbers for. But it's also the foundation for multiple generations of products. We've designed it from the ground up to support multiple generations of our products and in particular we designed it together with our next generation wafer chip that will be coming next year. And with that chip we'll be pushing the performance even further.
So we just saw a 2x improvement this this year with CS4. We're going to push it even further with another 2x improvement in performance. And what this ultimately means is you'll be able to run you know medium size models like GPT OSS or Gemma at speeds up to 10,000 TPS. And even frontier-level models like Kimmy or DeepSeek and GPT-5 6 all up to 5,000 TPS. Completely game-changing. Um all again enabled by this brand new platform that we've designed for multiple generations.
And so, you know, this is the trajectory that we're on right now and and you know, now that the world sees the value of ultra fast difference, our bet is that even this is just the beginning. >> What's the rollout process like? So, you have a deal with OpenAI over the next few years. Um when do we get access, you know? So, everyone wants ultra fast. Right now, sure there's Gemma OSS. When do we get the big Kimmys, the DeepSeeks?
When can we start to, you know, actually use >> a great question. So, right now we are basically, you know, sold out of everything that we're building, right? And we are very strategically making sure that we're deploying every single megawatt in the most strategic way possible. And in particular, quite a bit of it um is going to OpenAI, right? And we've been very public about this. You know, they're our biggest partner not just because of the the the commercial arrangement, but also because of the fact that we have this co-design um you know, spirit, which you guys have heard OpenAI talk about this a lot as well, right? which we believe will basically allow us to kind of continue this flywheel of not just making models, you know, faster, but using that speed to make them more intelligent and so on.
So, so today a lot of that capacity is going into OpenAI. And within OpenAI, they have very strategically decided to use quite a bit of it for themselves. Um so, they're using it right now. That's correct, right? So right now internally they're using it for a lot of really critical use cases where the speed really really matters. Like they're using it in like their incidents response teams, right? When there's an outage in their service for example, every single second, every single minute matters.
And so they're getting a tremendous value out of that. They're using it in some of the most critical research applications where the extra reasoning is enabling significantly, you know, more intelligent responses. And so right now that's quite a bit of the capacity is going there. And then as as we mentioned in in our launch with OpenAI, we are now also making it open to, you know, to enterprise customers who are, you know, who are who are able to again use it for some of the most demanding and, you know, the most high-value applications.
Over time, OpenAI and Cerebras, we have committed to, you know, to bringing enough capacity to make ultra-fast inference available to a much much wider audience. And you know, and our our CS-4 announcement, our CS-5 announcement, these are very much part of that commitment to continue to drive more faster tokens and, you know, basically more throughput to be able to satisfy, you know, all the various use cases out there. >> By the way, you know, the way that you framed it, I realized that they didn't have to expose it to an enterprise customers.
They could have just kept it for themselves. Nobody knew. That's an interesting like business decision almost. I don't know if you have any way in on that. That's a more like a business analyst point of view. >> Well, I will say I will say that, you know, obviously I can't, you know, explicitly comment about their thinking, right? Ultimately it's OpenAI's decision how they want to use this. But I think if you look at it from the outside, it kind of makes sense, right?
I mean, their mission very much is to continue to, you know, push what's possible, and by using it internally, that helps them do that. But, they're also business now, right? And there's, you know, talks about them going IPO and all this. And so, very much, even externally, you can see they're balancing both of these, right? And I think we very much see see that playing out in our um, you know, in in the ultra-fast space, as well. >> I mean, they're balancing a lot, you know.
Most recently, they're hot >> Got to talk about this, as well. >> We got to we got to go into them, so what are we thinking? Prefill Jalapeno, decode Cerebrus? What's the What are your picks? >> I think that's I think that's uh that's a very, you know, rational um, you know, conclusion. I I think of all of the hot chips uh announcements, probably Jalapeno was the most exciting um, but uh to me, but not maybe not for the same reason as everyone else. >> I guess to to double-click on that, you're the expert in chips, right?
There's people see token per second. What do you see as interesting when you see that? >> Yeah, so that's what I was going to say. I think that like they, you know, they pushed a lot on the performance, and the fact that they're significantly better than, you know, better performance than um than the GPU than than Nvidia. But, what I see is that they've built a uh significantly better GPU. And that in its own right is is very uh is a is a is a big achievement.
Nvidia knows what they're doing, right? They they they own the market for a reason. They're not dopes. Um, and so, to be able to come out of the gate and and build a significantly better GPU is a big achievement. But, to me, the reason why Jalapeno is so exciting isn't even these all these paredos. It's really um the design methodology behind it, right? They they very clearly took a very drastically different approach to building this chip, right?
Having an AI first methodology enabled them to, you know, build the chip faster and achieve some of these very impressive results, right? And that is 100% the future of our industry and it's not surprising to see that OpenAI is kind of leading the way here cuz this is, you know, very much their their MO. But coming back to, you know, what you said about you know, how we're going to use this, right? I think that, you know, Halapenio is is pushing the boundaries for both the boundaries of what's possible for both throughput as well as latency.
And for for me, even if I if I take the Cerebras hat off, I think that's awesome for the industry, right? This is going to lift all the tides. Everyone is going to benefit. Um, you know, they've they've also now been able to, you know, push the latency into regimes that that traditional GPUs can't can't hit. And and that's also great because again, we believe in speed. There's a lot of value there. And what's going to be really interesting and what I'm super excited about, right?
And you know, OpenAI is our biggest customer, so this is this is one of the things that, you know, I think is very strategically important to us is that when Halapenio is available next year, when our next generation CS-5 is available together, right? We will enable a full fast inference portfolio, right? That is substantially different and better than what's already available today, which is already substantially different than than your baseline GPU.
And then on top of that, there's opportunity to integrate even further, like prefill decode disaggregation, for example, or other forms of disaggregation. >> custom specialize Halapenio for, right? yeah, yeah. >> did not specialize Halapenio for that, but they specialized it for throughput, right? And so, they get a tremendous amount of throughput. And so, just as a computer architect, there's like so many different things you can start to do with that, right?
And then we have, you know, the you know, insane latency, right? Like I mentioned, up to 10,000 TPS in in CS5, you can start to imagine some some really, really cool products that we can build together, right? And that that's what really excites me about having them as a partner. And we're also excited about collaborations on the AI tooling front because much of the the benefits that, you know, that that they're seeing from the AI tooling infrastructure, um everybody probably says this now, but like we're obviously doing, you know, a lot with AI, but we're also collaborating very closely with OpenAI, right?
To use their tools to help us also continue to push what's possible in our chip design and our software and all that. So, both of those together, I feel like, you know, it's a really unbeatable combination. >> I think one thing that people are talking about that, you know, I I'm trying to get to the disagreements of the hot takes now. So, uh people are focusing on you haven't really mentioned like power. And I I do think that something that seems to be a consistent theme is um you know, performance per watt rather than than tokens per second.
Any variation of of this theme or or people sort of talking about offstage that, you know, is more contentious? >> Well, I think there's a there's a few things that you know, in terms of, you know, some of the the the more contentious things. Like I I would say one of the the themes that came up quite a bit in my discussions at the conference was around the the Groq announcement, right? You know, Groq, obviously not surprising, they have a new chip.
In general, I think it's it's awesome that yes, RAM designs are becoming, you know, more mainstream now. >> You've been here the whole time. >> We [laughter] we've been talking about it for a long time, and it's amazing to see, you know, the the industry starting to embrace it, right? Um >> Do they feel different post-acquisition? I mean, you've been competing with them for a while. >> to first order, no. I think it's a really awesome to see that, you know, SRAM architectures are becoming, you know, more accessible, you know, even the biggest of the big guys here, Nvidia is embracing SRAM design, acknowledging that you know, that uh the traditional GPU designs really can't hit the ultra-fast, you know, regimes.
A lot of what was being discussed uh offstage was I mean, the natural obvious questions is like, "Why didn't they launch on a 30 billion parameter model?" Um and, you know, how come when uh Jensen spent so much time at GTC talking about uh attention FFN disaggregation, there was no mention of that. And so, I think there's, you know, I think that's that's pretty telling, right? Um >> I mean, this is separate Rubin, you know, strategy. >> Well, so there's there's there's Rubin, but, you know, the LPX itself >> That that was supposed to be where it was? >> It was supposed to be Rubin LPX together, right?
Um if if you guys recall, I mean, the the Jensen spent like >> GTC, yeah. >> GTC like half an hour explaining attention runs here, and [laughter] and, you know, and and the MOEs run here, and so on. And I don't know if this is like a a a a hot take per per se, but you know, it it it it's it's very suspicious that their um that their product that's in full production, they've only shown performance numbers on a non-disaggregated 31 billion parameter model.
Right? And I think to me what this this shows is that there's definitely some challenges in running on a non-wafer scale SRAM design, because there's not enough memory in each of the chips, right? And I think that's you know, that that's what's happening and and and we're seeing the evidence of that. And in many ways, I think it's very much validating kind of the design choice that that that we had. Right? Uh if you think about it, to run a a frontier level model that let's say a few trillion parameters, you need thousands and thousands of Groq LPUs just to hold the the weights. >> Yeah. >> Right?
And so when you start to think about it that way, it's like, well, is it surprising that the only performance numbers that they're showing are you know, on 30B, right? >> So they're going to do gradient descent to this. >> [laughter] >> Well, I I think it I think it's I think it's the other way, right? I think what's what's going to end up happening is they're going to end up focusing on significantly smaller models. >> All right. >> All right?
Um you know, if you if you have that limitation in your architecture, then I I think that's that's what ends up ends up happening, right? >> Yeah. >> Um whereas in our case, you know, we're running the world's largest models and in some ways kind of simple cuz you know, one of our chips has you know, order 100 times more memory than one of their chips, right? So we got two orders of magnitude difference in scale kind of for free, right?
And so I think that's definitely one of the things that was a topic of discussion again, independent of of Cerebras, um just kind of odd that, you know, you'd launch a brand new product on on performance numbers of such a small model. But when you peel it back a little bit, it it kind of makes sense. I mean, I've been living in this space now for a long time. There's a reason why we needed the wafer scale integration of to be able to aggregate enough SRAM to be able to actually make it useful for large models. >> Yeah.
And look, the market's large, right? You have a different market than them. And you know, that you you you clearly are the longest-running incumbent now in this space. >> No, absolutely. I mean, the the the large, there's a lot of different opportunities for, you know, different different hardware to to play different roles. In fact, in general, uh you know, at Cerebras, we we we believe very very strongly in a heterogeneous disaggregated, you know, uh ecosystem, right?
Not just prefilled decode disaggregation, but you know, I think we're just at the beginning of what's possible in this space. And with the scale of of of these deployments and of these right, inference is no longer just like one workload. There's many many kind of sub workloads within it. And, you know, you really want to use the right you know, tool for the for the problem. You really want to use the right hardware for the problem.
And so, absolutely, I think there's there's a there's a spot for all the different types of architectures out there. And, you know, we we've we've chosen to to target, you know, the frontier. Um yeah, frontier ultra fast. >> Exactly. >> Um I I was going to go into some of the other companies that were, you know, talked about in town this year. >> But, it gives me This gives me an opportunity to follow up on one thing, which is how do you think about the classes of workloads?
Right? Um to me, the ultra fast frontier workload is just uh like very clearly like one of the fastest growing segments of the entire inference market. Right? Like I the the fact that like I want it and I can't pay for it enough for it. Um I might have been a mistake for OpenAI to offer it even, right? Like but it's good for Cerebras. Anyway, so I just in your experience, right? You say you say you're a strong believers in a heterogeneous um inference solution.
Okay, like what are the buckets? And um how do how do you see it from your talking to your customers? >> Well, so I I I think there's probably two different views here. Um you know, one is the the product view and then the other is kind of the the the technical computer architect's view, right? From the product view, it's actually just very simple. It's that bringing more speed opens up significant opportunities applications of applications different use cases and different capabilities more intelligent models more intelligent agents and so the further you can push that the more and more that you can enable in fact there's probably all sorts of things that we can't even imagine that you can build you know when you're even faster than what we are calling ultra fast today right which is why we keep pushing that and then the other angle really just comes down to capacity right that's so it's like yes I want the fastest but then you need enough to actually be able to to to serve your use case right and so it really just is that simple right and then finding the right architecture and the right mix of architectures right to be able to provide that is is really the name of the game now from the kind of computer architects point of view in a lot of ways it's it's even a little bit simpler right like I was having a conversation with somebody actually at hot chips about this and they were asking me well you know if you're doing disaggregation and like doesn't that mean you have to like partition the data center and you have to deploy a certain amount of this kind of hardware and you know different type of a certain amount of another type of hardware isn't that restrictive and I'm like I mean you know when we design a chip is modular and when we design a chip every single day we're deciding like am I going to use the silicon real estate for memory or if I'm going to use it for computer or if I'm going to use it for IO um and so from a computer architect's standpoint what I see right is there's significantly different parts of the workload right the simplest ones are okay just prefill you know decode but then you go one level deeper it's like well what's actually happening during these phases oh well there's you know there's loading the actual KV cache there's doing the actual you know attention there's spreading the experts across the various different hardwares and balancing them all of these things can be now addressed with kind of more tailor-made, you know, hardware solutions and and architectures.
And at small scale, it doesn't really make sense to bring together a bunch of different hardwares just to solve like one problem. But at the scale we're all talking about, the hundreds of megawatts to gigawatts to multi-gigawatt scale, it easily pays off. And then at that point, it's as if you're like thinking about the entire data center as if it's like one computer. It's like one chip that you're trying to figure out, okay, well, I want this amount of this capability so I can run, you know, attention fast.
I want this amount of capability so I can run um prefill really, really efficiently. And being able to piece all these things together is like the computer architect's like dream to have all of these tools in our toolbox, right? So, that's how I how I think, you know, where you get the value from this heterogeneous disaggregated ecosystem. >> How much of that plays into model design, model architecture? So, working with OpenAI, you get to hopefully see what's going on on models.
Like there was the whole MoEs were a thing, reasoning models. >> It's not enough. There's also voice diffusion. Uh what what if they what if there's more interactive models like the thinking machine stuff? >> Uh absolutely. And I think as as we start to get into more interactive use cases that we enable through UltraFast, we're even seeing a lot of this, right? So, if you're trying to interactively, for example, do graphic design.
Now, you mentioned diffusion. Now, you know, image generation or video generation might be kind of now inline in your, you know, creative like inflow uh interactive uh use case for with with the thing. And so, there's so many of these different use cases that are coming. And what I, you know, you mentioned about co-design with OpenAI and so on. What what I think is um the most untapped opportunity right now, uh frankly, for Cerebras, but frankly, for the entire non-Nvidia environment, right?
And the non-environment Sorry, a community is that, you know, we're all running models that were designed for Nvidia GPUs. Right? And And And it's not just Nvidia GPUs, but like usually they're designed for like one particular Nvidia GPU, right? Like, okay, this thing was designed to run on B200, GB200, MB172. Right? And, you know, and so at Cerebras, for example, here we were showing that you're running the model like 14 times faster already, but it was a model that was designed for a completely different architecture.
And so, if you start to then open up the possibility of adjusting that model architecture, even slightly, you can get massive gains. And if you start to do even more co-design, I think, you know, the the the opportunities are almost limitless. And that's really what excites me a lot about, you know, about working closely with customers and partners like Opening Eye. >> Yeah. And we We don't have time for this cuz we want to move on, but the people under really sleeping on AI code gen for kernels, uh which makes actually which is very good for you guys. >> [laughter] >> Yeah.
No, absolutely. I'm I'm a huge believer in that, for sure. >> I think before getting into etched and super spicy chips, any takes on AMD, Nvidia, what they're doing, could they be doing anything different, any hot takes there? >> I think that Nvidia is doing exactly what we all expected Nvidia to be doing, right? Like I said, they're no dopes, they know exactly what they're doing, and they're continuing to push throughput.
And in many ways, it's absolutely the right thing, right? Cuz we need more tokens, we need them cheaper, that's a that's a reality. However, you know, what used to be considered fast at like 1,000 100 or or 200 tokens per second is quickly becoming the new batch mode. Right, quickly becoming the new like overnight um >> It is true. >> And and so I think that you know, there's a place for that. It's great for prompt processing.
It's great for very very parallel workloads. Um >> But my eval stake 20 hours, I can hit a button and run it in 2 hours. >> Exactly. Exactly. [laughter] Right? But but quickly, you know, being able to do prompt processing isn't isn't quite enough. >> That's a big part of the market. >> But that's not quite enough. And so I think all of the the the traditional architectures I feel like are very much all on that treadmill, right?
In a good way, right? Not not necessarily bad way. It's in a good way, right? And so I AMD, Trainium, in many ways TPU, like all of these in my mind are all trying to build a better Ruben. >> [gasps] >> Right? And there's a huge amount of value in that. But I think that there's also a lot of value in trying to push, you know, the boundaries in other vectors as well, which is obviously, you know, what we're trying to do here at Cerebras. >> Okay.
Um spicy spicy chips, there's this page that you have on your website that uh this company Edge just doesn't have. No No numbers, no benchmarks. Any Any takes? That's That's our one criticism. But what's Edge? >> Well, so they have said very little about what they're doing. >> Mhm. >> Uh but what they have said has changed also a lot compared to what they originally said. Um >> Which they acknowledge. >> Which they acknowledge.
And and you know, the the environment is changing a lot, and so that makes sense, right? Peop- They're going to pivot. They're going to try to find their their space. I think my reaction when I see pictures like this is that it's very impressive graphics design, but I also don't see them building anything beyond just you're trying to build something better than just, you know, a traditional GPU, right? You know, they've made claims about having a distributed storage that is somehow faster because it's distributed, which I don't quite understand.
You know, it'd be very helpful if they can explain those kind of things. Um I think you know, I it's hard to see where the the differentiation is that they're trying to push. I think they had they had a rack or something at Hot Chips, I think, and so clearly they're they're building something, but it feels like it's a little bit, you know, far off for now. >> Yeah, but at the same time, like $20 million now, but you that's like, you know, not not a pushover.
Um I I think maybe the the steel man for this is like what interesting direction would you want to see them pursue, right? That In in the best case. >> In the best case, I would love to see them or others, frankly, push the boundaries of more innovative solutions outside of just the chip, right? You know, we I really believe that the next kind of frontier of of taking AI compute to the next level is all about better ways to integrate outside the chip, >> Right?
The The rack, the node, the the chip. >> the node, the package, the right? The the the interconnect. >> The factory. >> The the the factory. Um They may be doing something like that. It's hard to tell, right? Um but I think that this is not just a a point about Cerebras. I think for the entire industry, right? Um even if I I take my Cerebras hat off for a second, right? What excites me the most is our ideas and and creative thinking outside of the single chip.
We all know that the applications, the problems, the models are now so large that you can't just do anything on a single chip. So, everything comes down to the integration. Everything comes down to how do you bring together more compute? How do you bring together more memory? How do you bring together more interconnect? Right? And so, you know, >> Scheduling, pipelining, all those algorithms. >> All all of those things, right?
And And so, you know, of the of the you know, the reveals yesterday, I think probably one of the the the the more interesting ones to me I mean not not it wasn't new news, but was was D matrix, right? The fact that they're really leaning into uh into into DRAM, you know, 3D DRAM packaging. And like this is the kind of stuff that I think we as an industry need to be doing much much more of. >> Yeah, I was going to say that your backpack stuff reminds me of what the memory people are doing.
Is that analogy that >> There's a little bit of of that, right? Because you know, what we're doing with our backpack and and power delivery is very much, you know, a a very compact 3D package, right? Where we can bring power directly to to the face of our our chip of our wafer, right? And in in many ways, like you know, the the DRAM uh 3D stacking integration is doing something similar, but not with power, but now with with memory, right?
Um yesterday, we also announced that you know, we also uh started 2 years ago our um DRAM stacking program for the very same same reason, right? And cuz I think that figuring out ways to to do this kind of integration is the key to to unlock the next major step forward, right? And so, I don't know exactly what Groq is doing Oh, sorry. What uh Etched is doing underneath the covers, but you know, when you ask like what is is stuff that excites me the most about what we're seeing in the industry.
It's really more creative you know, integration, more creative packaging, more creative outside of the chip thinking. >> Yeah. Uh let's leave some hints for people other than D-Matrix a couple other names that stood out to you. Just >> Along these lines I I think you know, Samsung has some discussions around their Z HBM. Right? Um >> Yeah, all memory guys. >> [laughter] >> I mean Look, in the end there's only three main things to to building a you know, a an AI chip, right?
There's the compute, there's the interconnect, and then there's memory. Right? And so um and and more and more importantly as we all know for at least in the low latency space, the memory is the key, right? The memory bandwidth is the key. >> power and cooling not as major still. >> No, absolutely. I mean >> I mean you you you just did the redesign. >> Tho- those are those are what's powering all of this, right? Yeah. >> I'm just saying you know, there's physical limits. >> Well, that's actually a really really good point, right?
I think that um when we when we originally started with wafer scale, most people thought the biggest challenges were okay, how do you connect all of these die together? How are you going to yield this thing? And those were absolutely kind of fundamental challenges that we had to solve, but what turned out to be you know, the biggest enablers was ultimately how do you power it? How do you cool it? Right? How do you build it reliably at scale?
And I think that a lot of the the 3D DRAM technologies are going to go through the same thing. And it's you know, one thing to draw you know, a PowerPoint slide with a DRAM and and a and logic, and then it's another to actually make it work, to figure out how you actually going to power it, how you actually cool it. And that's why a a lot of what we've been doing is you know, is trying to leverage and use all of the expertise that we built up there and put it into our DRAM stacking approach.
And so, you know, we already have solved yield at scale, for example. We've already solved how you can actually package in a three-dimensional way. And and that's, you know, problems that Samsung, that D-Matrix, and and everybody else are also going to have to solve over this time. >> Thank you for those comments. Uh last topic because we got to go. >> Okay. >> Uh US supply chain semiconductor supply chain stuff, you know, like uh China is slowly becoming completely independent of us.
Recently, they uh >> They are. >> They are. They are, right? >> like we literally just like they're they're they're bragging. And they're totally Huawei and what what have you. >> Hyped outside of knowing what it was on. Just model is good, doing good in, you know, open router. Came out and not even running on US chips. >> So, the theory is are people talking about it? Uh what's what what are we doing? >> I mean, I think that absolutely, I mean, of course that people are talking about it, right?
Cuz like the the the we're in a very uh difficult, you know, situation right now. The open source model market is 100% Chinese, right? Um 100% but almost. Okay. 95%, right? Most of the most of the big models, most of the big open models that are, you know, call high quality are uh coming from the Chinese labs. In some ways, it's great that sharing is happening and and, you know, the the global community is benefiting.
But also obviously, that's a very strategically challenging place to be, to have such dependence on on them on the model. Separately, as is evidenced by this, is like, you know, we have independently the all of the infrastructure, the hardware infrastructure, slowly being built up in the background to support all these Chinese models, right? I think it's absolutely a a problem that uh globally as well as as you know the the the US and as a as a natural interest that we have to continue to push the boundaries so that you know we can continue to to compete and and continue to be ahead.
Um it's not easy. These guys know what they're doing. Uh but it's something that I believe is not solved by you know one company or or one fab but it's needs government level national interest level kind of initiative to be able to make this happen. Um >> So I I imagine that should take place at at Hot Chips. It's like a secret room of all you guys. You're leaders of our industry. Right? Like you would be the guys to >> I mean we we have absolutely been been pushing for this and we've been very supportive of of you know any initiatives that will um you know continue the US dominance in this space 100%. >> Okay, uh you got to go but you've been very generous with your time.
Congrats on all your success. I mean enormous you IPO'd. >> [laughter] >> Yeah. Yeah. Yeah, sometimes I still have to pinch myself to remind me that we IPO'd and it was like the largest like you know semiconductor IPO in history and it's like Yeah, that was >> early. >> That was not bad. That was not bad. >> Uh no, congrats and uh look forward to meeting up with you in future year. >> We got to ask, you know, give us ultra fast, give us bigger, give us better models. >> [laughter] >> No, no, I mean give give the people ultra fast too.
Like But we but we can probably get you in line at at Open AI. >> I hope you didn't buy us Open AI too much to you know just keep it internal you know. People need it. We need better models. >> It's it's Open AI's decisions. We're we're just providing the infrastructure. >> Hey, but give us give us GLM, give us Deep Speed. >> Just give us our own rack and then we'll run it. Uh okay, well thank you very much. >> Thank [clears throat] you again for coming.
Really really appreciate it. >> Yeah. >> [music]
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.