Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
Most replayed moment at 13:49
2.3x that video's typical replay level
Fable uh and it runs into an unknown, ask it to log it, right? So that um you uh you can see where the deviations happened and then you can sort of figure out why as well, you know? It will usually give you some context about what happened.
Said at 13:43
The graph counts replays. It does not show where viewers stopped watching.
Words
8,585
Runtime
43:20
Speaking pace
198wpm
Reading time
36min
198 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
[music] >> So, I hope everybody had a great lunch and you got to check out some of the amazing demos that we have. Uh we're going to begin the panel the first panel of the afternoon here where we're going to be talking about of course the engines that are actually powering the stuff that you know could remotely be used for things like local sovereign any kind of ownership over your own artificial intelligence and of course the engine powering those in addition to the hardware is the models themselves. And so for this panel we have
99 words, the words spoken in the first 30 seconds at 198 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 299 |
| Average words per sentence | 28.7 |
| Longest sentence | 111 words |
| Questions asked | 44 |
| Sentences containing a number | 21 |
Most used terms
Filler phrases
691 in total: like 277 · uh 142 · um 85 · you know 85 · right? 33 · kind of 29 · actually 13 · sort of 13 · basically 9 · I mean 4 · literally 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
[music] >> So, I hope everybody had a great lunch and you got to check out some of the amazing demos that we have. Uh we're going to begin the panel the first panel of the afternoon here where we're going to be talking about of course the engines that are actually powering the stuff that you know could remotely be used for things like local sovereign any kind of ownership over your own artificial intelligence and of course the engine powering those in addition to the hardware is the models themselves.
And so for this panel we have excellent guests. We have Vincent who's the CEO and founder of Prime Intellect. We've got Lucas the CTO of RCAI and we've got Chris who is the senior product research engineer on the Neumotron family of models at Nvidia. Now, what's really cool about working in this industry is really cool companies like this we all get to work together. And so this is one panel where we all directly get to work together on both models infrastructure some of the ways that we think that the direction of the industry should go and each of us kind of play a different role in that stack but I want to leave it to you guys to to introduce yourself and be able to talk about sort of the the charter that you see the problem of the stack that you guys are working on. >> Awesome.
Should I kick it off? >> Kick it off. >> Yeah, so I'm I'm Vincent as you mentioned and then really the goal with Prime Intellect from the beginning was like to ensure that basically frontier intelligence will be open and accessible not just the models but also the full stack to to train the models. So kind of like this was like our motivation from the beginning and we we've yeah like worked also together with a lot of gentlemen on the on stage like on the one side is like we we work with folks like Lucas and RC to help them train frontier open models.
We we we help also um like Nvidia on the Neumotron coalition help their uh open models. And I think like I'm actually um think both like an Eleutheron and and Trinity might be the best like two open models right now outside of China. So, I think it's actually uh like we need to fact-check that, you know, but this is actually from my I think they they might be. >> It's our our marketing says they yeah. >> But yeah, so so that's the high level. >> Um my name's Lucas Atkins.
I'm happy to be here and and thank you for joining. Um very similar to Vincent, RC was uh you know, founded with the idea of um domain-specific owned models are are are going to be needed. Um you know, we were founded early 2023 uh jumping on the custom model uh train quite early. Um you know, you have all these people who are excited about AI and all the things these new generation of LLMs can do. Um but they're using these monolithic very expensive closed APIs uh for at the time and still like very narrow tasks that don't require uh you know, at the time it was $100 per million tokens out.
Um and through doing that, we were building on top of open models and we were releasing a lot of our tooling in the open. Um and uh we noticed that in the United States um and in the West, you know, in general, we were starting to lose uh leadership in the open model space. A lot of it was coming out of China and that's amazing. I love those models. We learn a lot from them. Uh we're close with a lot of the people building those.
But when you're working with large enterprises and companies and um geopolitics gets involved, whether you like it or not, you have people that become concerned about where those models are coming from. And we uh decided that, you know, we had a a good group of people and we had a good group of partners like Nvidia and like Prime Intellect where we could probably try to pre-train ourselves. Uh so, last year we did that.
We kind of uh reoriented the entire company towards let's figure out how to pre-train a you know a 400 billion parameter model in 6 months. And a lot of people said it was impossible and by in many ways it was, but we figured it out and now we are an open model lab working with our wonderful partners and our customers to build western open models that are permissive and you can own those and customize them or run them wherever you want and and that's kind of where we're at right now.
So thanks for having me. >> Yeah. >> And so I'm Chris Alexa. I work at Nvidia as a product research engineer and I support the Neumotron family of models. I think it's you know we we've talked a lot about why we do Neumotron but just to say it a few more times. You know AI should be open open as in weights, data, training methodology, training frameworks. Really respect a lot of the work that uh the the the the two other people up here do because they they believe that very strongly as well.
But the Neumotron family of models is focused on being as open as humanly possible. So we we have this understanding or belief that in order for AI to continue to grow and be useful to everybody, it has to be done in the open so that we can build off of each other, we can compound on each other. And part of what we do because team green, this is always true, is we we think that the the rate that you can squeeze tokens out of models is very important.
So we kind of have this mantra that like faster models are smarter models and so a lot of the decisions we make when designing a model like Neumotron is built around how fast can we make it go. As especially you are going to see in the next however many months local AI take off, we we need to make sure that models are well supported on uh hardware that doesn't just exist in massive buildings, you know, thousands of kilometers away from you.
Uh and so that's uh you know, for AI to be very useful, it should be quick uh and and open. So, that's kind of the the vibe of Neumotron. >> Who makes those buildings with the massive processors? >> Oh, that's uh a lot of excellent people in the world that they use a lot of excellent hardware from a pretty cool company. Yeah, I heard I heard anyway. >> Yeah. And I'm Carter Abdallah, I'll be your moderator for today. Uh something that you know, we all kind of talked about is is this you know, building on top of each other, learning from others, whether it is people, you know, across the the big pond of the Pacific Ocean from us.
Um but really it is kind of like a collaborative sort of research effort and I imagine that a lot of the people here in this room share that sentiment. But as it was brought up during the, you know, inaugural panel this morning in the state of the union, there there is a growing sentiment um potentially on the other side that that paints uh open source to be something that is actually more chaotic, that there there's less trust involved.
And I think trust ultimately as Lucas, you and I were talking about before, um depending on who the party is and depending on what lens you're looking at at it from, I think it kind of means different things. But ultimately from from the the end consumer, the somebody who's using this intelligence, um or somebody who's you know, more of a business and is actually customizing something to maybe monetize tokens in the in their business.
Um can you comment, we'll start with you, Lucas, a bit on how open source and open source models are actually key to building that trust so that when these people walk out of this room and somebody does come at them with that other angle, they can they can sort of steel man this side. >> Certainly. You can weaponize any term uh and certainly trust is is has been weaponized, that word. Um and the reason I say that is because it means something in based on the context and with your speech you're speaking about it.
Uh often in AI, people like to conflate trust with safety um and those are not the same thing. Uh, and I'm happy to speak on safety, uh, you know, later on. But when it comes to trust, um, I I think that, you know, you hear a lot from closed model providers or uh, politicians or people out in the space who are advocates for or uh, against open source that you can't trust these open models cuz you don't know what went into them.
Well, the same is true for these closed models, uh, even more so. Uh, the benefit of uh, uh, of of open models is that uh, we can very easily validate what is inside of them. Uh, they are you can there is a whole bunch of files with a whole bunch of matrices in there and you can view them and you can see the code that is running these models. You have implementations from Prime RL, vLLM, SG Lang, the provider themselves.
These models are inherently trustworthy. You know much more about what's going on when you hit and talk to these models than you ever will what's going on when you hit an arbitrary API. Now, that being said, um, certainly there is fear that people can um, reduce uh, you know, the you can't trust that these models are writing safe code. Well, again, that is the same thing with any model. You need to you need to use your uh, your judgment and you need to make sure that you have the proper um, you know, safeguards in place and you're viewing the outputs of these models as the outputs of uh, an inherently random system that we are working very, very hard to make less and less random.
I think that uh, a telling thing is a lot of people said, "Well, you can't trust Chinese models. You can't trust Chinese models. You can't trust Chinese models." That was often uh, for the last few years meant you can't trust open models. Well, as soon as uh, Anthropic had to put Fable away and people realized that, "Oh, our access to these frontier systems might not be universal anymore. Uh, there's probably going to be a lot of checks and balances.
You had a tremendous number of enterprises and developers and companies start going to these new Chinese models because they could trust that they would always have access to them. Um and so when when it comes out of trust in the way I view that word as it relates to open models is do I know that what I am running and can I be as sure as possible that when I send something to this model that I am going to get the output that I expect?
Um and the only way uh currently to be 100% sure that what you are getting is what you were expecting is by hitting an open model either that you are running yourselves or you're working with a partner like Prim Indelec or RC or Nvidia to validate. So that's my take on the word trust. >> I think too something you mentioned is like we don't get to know a lot about the data that goes into these models and that's something that I'm really happy you know that that we're trying to do which is not it's not you know the incentives don't exist for everyone to do this right so it's not something that I think is mandatory or should be mandatory thanks to the things that Lucas mentioned which is that it's rather straightforward to validate what data did go into a model without seeing the data sources originally but I'm happy that Nvidia continues to release data sets along with our models.
Release environments along with our models to make sure that even even if you can't go through the work of determining what went into the model which you can do with the the weights alone for the most part you have like a spreadsheet you can look at that says here's you know a couple trillion tokens of this data set a couple trillion tokens here and I think that that helps to educate people on why it's much easier to trust open models than than models that we we don't get access to any of that. >> helps people see what that data looks like. >> Yeah. >> You know if you don't have someone releasing it openly when someone says data is going in I mean data can take many different shapes.
You can but you can go to Hugging Face you can go to Nvidia or you can go to Prime Intellect, or RC's uh HuggingFace. You can look under our data sets, and you can see exactly what that looks like, um and that can help you uh inform your priors on it. >> Yeah, I think that trust also um you know, there's some angle of of a reputation. Do I believe that your your intentions are pure? And I think that a lot of people, again, in this room believe that intelligence is kind of this next layer of almost, you know, infrastructure for for us to progress as a species, and I believe that everybody should have intelligence.
So, uh on on that front, I want to uh hand it over to Vincent because uh you kind of have this um almost like founding thesis that you this this stack should be the open, right? The open superintelligence stack. You want everybody to have a lot more intelligence. Um can you talk about how uh the this is kind of moving into the era of control, but uh beyond just data sets, how important it is to have the knobs and dials of the of this industry also be uh available in an open way for people who are building this? >> Yeah, like I I think it's a really important point is to set um so I can say it like be able to take those open models, like customize them, be able to like build on top of them.
And I think like all the different components are going to it, like especially from like the pre-training to mid-training to post-training, I think >> [music] >> like need to be more accessible, right? Like so more people can also like take those uh amazing models and like make them work for their specific use cases. So, I think when we started like we we also took a look at the whole stack that was out there and and tried to figure out like what is missing for ourselves to train open models and for like helping our partners to do so.
And a lot of this was around the RL and post-training stack. So, we basically went deep into building out like a lot of infra around that, like around our environment, Evals, around like making it much more accessible to do post-training also because it's like the most economically viable way to maybe like customize those models, to take an open model, and um to have like a specific Eval environment, and the um specific domain and dimension that you want to improve it on.
And this This of like what we're are doing with Prime Intel now, is like enabling people to post train specialized agentic models. Um so, being able to take models like um Trinity, for example, from RC or Nematron or others and and specialize them, post train them for the use cases that ultimately enterprises care about. So, good example of this was like a company like, for example, Ramp or Saber to like take an open model and like specialize it to automate finance within like a week or two to get like better performance than like Opus at a fraction of the cost of Haiku.
And I think really this Pareto frontier of like being able to create these specialized models that are much better than the frontier, but also faster, cheaper. Um I think it's like a key thing enterprises care about increasingly. It's really like just making it work for their use cases, basically. >> If you go back to trust, it's how you can make your CFO trust you by knowing exactly how much something's going to cost all the time.
Um that's That is increasingly becoming very important is uh you hear a lot, you know, all these companies have unbelievably large token spend and um they're having to cut back on their Opus usage because they burned through it all in a couple months. Um and that is going to continue to be a problem because yes, the cost of an individual token has come down drastically. You can look at it, you know, the difference between GPT-4 when it first launched and GPT-5.5 is is is much, much cheaper per token, but at the same time the amount of tokens in an individual session has gone up exponentially as well.
So, we're kind of um we're spending more uh on a on a total session. And so, the ability to uh bring in-house or or or at least work with partners to ensure that you are controlling your cost and you're not at the whims of uh when a company releases a newer model uh that might be better, but also more expensive. They might deprecate a model. Um owning that and being sure that, you know, same way is what you what out input goes in, you know what output's going to come out.
In the same way, um when it you know, an input goes in, how much it's going to cost. Uh having assurance on that's is important, too. >> Yeah, and maybe like one thing to add to this is like almost like I like this new term of like instead of speaking about token maxing, you know, speaking more about like the outcome maxing of like >> Yeah. >> Ultimately, it's like you you want to have like more than a dollar worth of value come out of like a dollar of input and I think this is sort of like Jevons paradox of like if you can create more value for like your flop for your GPU, basically.
I think this is sort of like how you get like the most adoption also of like agentic models. Like if they can like be able to create as much value as possible. I think the cheaper those models get, the more usage they'll get like for those specific use cases. >> It's funny you say that. I have a and I think a lot of us in this room, but especially on this panel believe this to be so that the most meaningful AI applications the next couple years, even this year, are the ones where the harness and the model and the product, they all kind of blend together.
If you think back to at least for me, the first like truly game-changing agentic experience I had was when Deep Research from OpenAI. And that was because they spent a tremendous amount of time doing reinforcement learning on O3 with test time compute to do these longer running research tasks that people had tried previously, but they were kind of just doing a for loop over search. Whereas I kept coming back to to Deep Research.
And you know, you saw for a very long time that OpenAI and Anthropic and Google, when they'd release a new product, they'd release a custom version of their model for that product. And if they're doing that, if their off-the-shelf GPT-5 isn't good enough for, you know, their Atlas web browser, why should it be good enough for our apps? And that's why I appreciate the work that, you know, Vincent and Nvidia are doing for for giving people the tools to customize their own model and allowing us to focus on how we get a good model to start from.
So it really is you know, it's it's extremely important as you look at developing applications and experiences over over next few years that you're taking into account that you can uh make the model do something that maybe your harness isn't fully able to do alone. >> Well, that's something I think that's really important to just like reiterate, right? I mean, like Nemotron's great, I love it. Trinity's great, I love it.
Like we design a model that's supposed to be as good as it can be across a number of harnesses, right? You can see this in the technical report. The idea is like we want the the model to work as well as it can in Pi compared to you know Hermes compared whatever you're using, right? But like you're you're not using all of these tools at once. You're using one of these tools. And so when you have open models, you can do things like news research can create a post train of whatever model for their harness, right?
That you know will be extra good. And you know, this this this thing from from from you know, I I can't remember who who who originally wrote it, but this idea of like the mismanaged genius, right? We're we're leaving a lot of like a lot of important capability on the table because we're just not we're not fitting the models into the harness, right? You can do a bunch of stuff with closed models. Like you can change your prompts and your skills and all kinds of other nido things, right?
But nothing will let you get the the the level of customization or customizability that you can achieve with open models. And I think that is something that is going to become increasingly and increasingly more important, especially thanks to folks like the others on the panel where you know, I can just straight drop like my favorite coding and you know, an agent environment spin up the CLI and suddenly my my model feels way better with very little effort, right?
Like that is that is something that is already at our fingertips and it is only going to get easier and easier as as time goes on. >> Yeah. >> And it's maybe also the most concrete like info for the builders in the audience like call to action of like if you kind of want to build the next like cloud code, the next like cursor or perplexity, I think the easiest way to get started is like take the best open model like and and then post rate on your harness like that you care about, right?
Like basically like create a product that is like truly AI native, right? And I think this is sort of like I think one of the most exciting like unlocks for builders um like out here. >> That's a big thing about the the framing of control is is um similar to like, you know, when when the cloud um explosion started to happen in the in the, you know, the late uh 2000s, early 2010s and the um social media world kind of took off and apps became extremely popular and more, you know, cloud-based and and, you know, managed by these these bigger companies, uh you know, you the data that they were collecting from you, whether anonymized or not, was how they were monetizing their platform through ads or or or or in other ways.
Uh and a very similar thing has always been happening, but I think it's becoming clear to people in in the space is that the the data that people get from from you using these models is how these companies largely make their models better, whether it's through actually training on that data or um by using it as a signal for what data they they go out and find or generate to train. And uh with closed models, there are terms of service that keep you from being able to um well, I could get into a a discussion about what terms of service is an agreement between you and the provider, it's not a legal uh anyway, but you you you shouldn't be training on a, you know, a Claude Opus output or a Fable output or a GPT-5 output, and they do a lot to try to obfuscate to to make that not great for you.
If you're using an open model, you can save all of those traces. All of those traces of you using it inside of your harness uh that will allow you over time to if you say, "Hey, I want to go train a custom model," you can take all of that, and again, either use it to directly do like fine-tuning on a smaller model, so you're not spending as much, or to have a model help you find signals, so that you can go out and use Verifiers or Nemo RL or Nemo Gym to create these environments, so that you can hill climb and make your models better.
So, um as much as using open models is like owning your stack, owning your intelligence, it's also owning your outputs, right? Owning your data. That's going to be extremely important, too. >> I do want to I do want to plug the license for a second. So, >> [laughter] >> uh so, AI is very different than traditional software, uh which is why recently uh Nemotron, as well as uh Trinity, I know, uh has adopted the Open MDW model uh data weights uh license.
Uh the idea is like, we we need a way to really make it very clear in the license that you can use the outputs to produce a model. You can use the outputs to train uh right? All of these TOSs and stuff like that that that that that have language that's meant to dissuade you to do that. Uh we wanted to make sure there's a license that exists that not encourages you, but makes it crystal clear that it is it is permitted it is permissible.
Uh and and I think, you know, the licenses maturing, right? To fit the use case better should be extremely positive signal uh for for the way that the ecosystem is thinking about open models. Uh to to the fact where even even the lawyers are on board. >> know how much lawyers cost? >> [laughter] >> Yeah. >> Yeah, it's uh as we move on to the, you know, kind of the third topic which here is of course optimization, um something that you know, we've implicitly said, but haven't said it quite explicitly yet is uh effectively that I I think that for a lot of people there's this preconceived notion that when you're deciding to use an open model for whatever the use case, um there there are the tradeoffs that come in the form of performance at the benefit of getting things like, you know, maybe data sovereignty and so forth.
Um, but what we are now talking about is that with the with the right customization and optimization, depending on the the use case that you're and the harness that you're applying it to, you can actually exceed and build the model against the tool to to get better performance than even frontier models. Um, I'd love to hear a little bit more about the cuz I think that uh another thing that we would probably agree on is that the the current level of intelligence um, already has so much uh left to diffuse into society.
And so, where are those areas where that diffusion is is is happening in the specific industries? I know, for example, things around uh again, kind of fundamental pieces of infrastructure, whether it's like browser use um and how you can start to train models to to be able to use uh you know, the the internet better when looking at a computer and so on and so forth. So, what are the what are the some of those examples to where you think that the uh post-training of open models will uh see new use cases basically unlock compared to just paying full price for the the frontier models? >> Yeah, I think I like like I can start on this.
Like I think the the power what that we've seen with a lot of different customers is is really kind of this idea that like if you want to make a specific use case work, like we can take the example of like if you want to figure out like a way that agents can actually automate your text. Like the the the most like concrete way you can do it today really is like build an RL environment for that use case. Like train on it and then deploy it into production with those users, right?
Like let's say with a a million accountants that then now use this agent to ultimately get it towards full autonomy. It's a bit like almost like Tesla's levels towards full autonomy where like you kind of need to deploy it into like do the last mile of actually like training for that specific use case, but then also um deploying it to those specific users, right? So, like there's a reason why like uh chatbot isn't good at self-driving because like it's not trained on that.
It's not deployed into that context, right? And like I think it's the same even for these specific like knowledge work use cases, where it's like if you want to have the perfect like financial agent, it's much more likely that you'll be able to get there if you have like our environment for that use case, if you deploy it into production, for example, as a bank, right? Like to millions of customers, than if you're there's like one got model chatbot.
Like And I think this is sort of like what we've seen now with a lot of verticals and customers that like um there's like a huge unlock there to um really go into this like specialized domains, post train on them, deploy into them, and then continuously learn from production traces. So, we work with like some um also big AI natives on things like computer use, where ultimately having like millions of of traces from production data really can help you to to um continuously improve uh those those agents.
And I think this kind of applies to almost every single domain, and I think it's sort of the the white pill for like the AI application builders and and the AI startups to actually have a a huge opportunity to build kind of their modes and and to get to this data flywheel um of like specialized models even in a broader sense. Like just going after like let's say computer use agents, right? And I think um yeah, this is something where I think um we're just seeing a lot of like movement especially now with like all models catching up to the frontier.
Um and I think the other piece is like optimization, where like I think um like GM is a great example or like also like Trinity and NeMo Triton is like you have the whole ecosystem sort of like driving down the cost and optimizing it further, right? Like like we very heavily work like very closely with all the teams here, but then also um like deeply also with Nvidia and with teams like ViLM to really drive down uh the cost and and make the for example inference and training for models like GM or like models like Trinity and NeMo Triton extremely efficient, so you can basically drive down the cost like further and further.
And I think this is something you don't obviously get with the closed APIs, where like they have like a huge margin on top. Like they might drive down the optimization, but then might not pass through those savings. So, I think in general, like the open models are only getting through the open ecosystem like more and more efficient like and and cheaper and cheaper to run and train on. So, I think there's like this element as well. >> I think too like a couple things that I I want to make sure we're very clear about is you you like most people probably do not need frontier level intelligence for like 90% of their tasks, right?
Like like not to say that you're not doing cool smart stuff. Not to say that I'm I'm sitting here trying to do not cool smart stuff, but like a lot of the time these models are just overkill or they have like this really smooth you know capability horizon that means they're they're also quite good at chemistry, but like most people are using models to do one or two things very well. And open models let you choose those one or two things and then make the model just very good at those things at the expense of at the expense sorry of almost everything else.
And that that is great. I mean that's exactly what we should be doing, right? To to to use this model that is hyper generalized and able to you know perform well across like 90 different axes is is dope and cool, but it is not really you know using the model effectively. It makes sense for someone who is trying to ensure that everyone can use this one endpoint to do their task, but it makes much less sense when you're a person who's trying to do that task yourself.
What what what was just said about efficiency is also deeply true, right? I mean the idea that you are all here at a local AI summit. Presumably you are running AI locally. Presumably you would like it to be faster and better. And presumably many of you are quite uh, quite cracked engineers, right? This is a whole room of people who is going to contribute in some small part to making the ecosystem just a little bit faster, just a little bit more efficient.
And while it's true that closed companies can, uh, afford to hire great amazing teams of people, as we saw with Linux over the uh, whole time that it's existed, right? Uh, Linux is the thing that runs the internet, it runs networks, it runs all of these services that, uh, that require it to be hyper optimized in a way that I think you can only get when you have people who are trying to run as resource constrained as possible.
And, uh, all of that to wax poetic and say this idea that like local AI and open models in the most efficient version of the of the model ecosystem, uh, is is necessary to do it in the open. I think it's, in fact, not possible to do it behind closed doors cuz you're shutting too many people, uh, that could make that one small contribution, uh, out of the room. >> I I think that that it's important to state, too, that I don't think any of us agree that or or or of the mind that closed models or frontier, you know, what OpenAI and Anthropic, just to name names, you know, Anthropic and others are doing it is, uh, not extremely beneficial or that not don't use them.
Um, I I I certainly, uh, I use those models near every day. It's it's just that it's where does it fit in the, uh, in in the future of this ecosystem. Um, and just like, you know, Chris is alluding to, you can think of open models and self-hosted or kind of owned intelligence or LLMs as like the Linux layer, which you're beginning to see kind of take place. Linux runs enterprises, you know, it runs the cloud. We're seeing a very similar thing take place with hyperscalers and neo clouds and providers like Fireworks, Together, Base 10's models.
Um, but just like Macs are one of the best ways to get work done individually in the same way that maybe using open AI and chat GPT is the best way for you to do the vast majority of simple check my email, help me rewrite, you know, check for grammar, those kind of things. It's accessible, it's easy, and for you know, your average consumer, an individual, it's pretty cheap. Um, if you're using like the $20 a month plan.
In the same way that you know, Microsoft helps the the world of medium size to large businesses run on Microsoft and Windows because you know, they're not as expensive as getting everybody a Mac, and you're going to see a similar world play out there for for some closed and open open frontiers themselves. So, it's all an ecosystem, you know, I don't want to give the impression that you know, I think anytime you log into chat GPT or Claude that that you're committing a sin, only that as you are, you know, this is a a conference for AI builders, AI engineers, as you're looking at the best way to engineer your product or your service, that there is another layer you can go down into and it's becoming way more accessible than it used to be. >> I I think that's a great point, and I think that the relationship between closed frontier models and open models will be one that is it's constantly there, right?
I think that we have it's never you'll never get the headlines to apply nuance and say that both will coexist and gain more usage and are going to be useful to everybody, but that is kind of the de facto state that that will not only currently exist, but will continue to exist. Um, I want to spend the last few minutes here to really give the audience something that only you guys potentially can answer. Often times I reflect about my time at Nvidia and I think I I feel as though I have a clear vision outside into there's definitely still a fog of war out there, but I have a vantage point that many people don't have and you guys because of the positions you are in as well.
And so what is the the thing if we're looking forward towards AI engineer World Fair that you think if you were to make a a bold prediction, let's say, let's not be conservative around the intelligence in in the open source and ground it it was some frame of reference. What do what do you think we can look forward to by by this time next year? >> I think one key aspect obviously that people are closely tracking is like sort of like the just like capabilities of open frontier models and I think they'll they'll keep being very close to the general frontier potentially even like what like what like now the the speed or like the of those close frontier models like slowing down.
I think like they'll they'll catch up even more and I think the the most concrete thing that I think will be very exciting is like seeing the world move from sort of chatbots and now coding agents to like just general knowledge worker agents. I think over the 12 the next 12 months right like to see more and more of like kind of like everyone across every knowledge worker domain like adopt agents in their workflows which I think like developers have with coding agents have probably done better than any other domain in the world, but I think we'll we'll see over the next 12 months like a lot of like domain specific like knowledge worker agents, but then also I think domains like computer use agents and I think others will take off.
Like I think in a similar way that like coding agents have taken off. I think we'll just see almost like in some ways you could say like almost like the general intelligence for like the knowledge worker in digital domain before then hopefully maybe moving on to the physical and I think very concretely like I think we'll like in in 12 months I think it's pretty likely that we'll have like better than fable metals level capabilities and open models.
And I think this is like a huge opportunity uh to ultimately enable a huge crop of new like AI startups and companies. So, there's some ways like you want to almost like write the the levels of capabilities. Like to some extent like cursor really only took off when like Opus was good enough to do coding, right? So, it's like this is when like cursor inflected. And I think we'll see like hundreds of these inflections for like startups getting started like now, like in over the next year or two.
Um once like open model and and I think we've seen this literally a month ago with like GM 5.2. I think like uh it was I think legitimately one of those moments when people were like, "Okay, this is now like similar to the Opus inflection point. Feels like an inflection point for open models to be like extremely strong and ultimately enable a ton of new businesses. And I think this will only continue like >> Uh not so bold prediction is that uh Primin Electin RC are going to have a combined valuation of a trillion dollars.
Uh that's all obvious. Uh but but but I think that this is going to be a huge year for um uh this is probably going to be the most consequential year for like the future of how AI gets distributed. Um the the the Fable uh and GPT 5.6 um you know, uh uh uh embargo, if you will, um has left a lot of open questions, uh you know, no pun intended about open models and and where um how this intelligence gets distributed. And at what capability level it starts to be uh politicized and and and kept back.
And so, sovereign intelligence is going to be very important. I I think that um if you were to take, you know, the general population of AI um users, people that are using it every day. So, you know, upwards of uh a billion to two billion people if you, you know, include ChatGPT and Google and whatnot. Uh maybe 0.000001% have ever used an open model, You know, I think it it or you know, run it on themselves with with the multitude of different tools.
Um and I would hope that with the work that we're doing and the community's doing and the way that we're advocating um for open science and open models and open discussion, right? That's probably been the most frustrating thing about the last couple weeks is that all of these conversations around capabilities and who gets to use them and who doesn't have been happening behind closed doors. Uh my hope a-a-a-a-and I hope that I can predict that we will be able to have a a 10% to 15% of people that have ever used AI have used a model locally on their system and that that becomes a very important part uh of ensuring that you have access to what you need.
Um and and so it's a prediction. Uh it's also something that I know all of us uh up here and and you out there are going to try to fulfill and I hope that uh we can continue to advocate for that because if we're if we're quiet, if we just let this uh things play out the way they are, uh open models will, you know, will will be put under the microscope um in the context of untrustworthy, unsafe. Uh and as much as there's work and and um vitriol and weaponized terms being out there uh advocating for that, we need to be uh combating uh as much of that if not more uh with the reasons that it deserves to exist. >> Yeah, I just uh I could not uh plus infinity what the last part what Lucas said more.
I think this is going to be the most consequential year for uh open intelligence that uh will a-a-a-a-at least for from where I sit determine the future of uh of a summit like this, right? Uh I think it will has a potential to look very different in two radically opposed ways. Uh as for bold predictions, I think that we will not be needing to go to an API for most of the tasks that we all do each day with AI. I think it's likely to assume that you'll be running a model that is sufficiently capable in let's call it day-to-day work on your on your MacBook within the year.
It's already extraordinarily close, so not maybe not that bold of a prediction to be honest with you. I also think that we're going to continue to see models become the the future of AI, so not model, right? Swarms of or specialized systems of models, I think you're going to be increasingly important. And lastly on the open model front, I think we're going to see some very large architecture shifts, especially as we start to crack things like diffusion models for text a little bit more to to get us models that are better suited for the the hardware that we have in our houses.
And then last meme one, I think you're going to buy computers with agent operating systems on them instead of traditional operating systems. Similar to like you buying a Spark preloaded with with Hermes or whatever. I think that's that's likely to occur. >> I also predict that come September when the next iPhone comes out, you're going to get a lot of texts from family members asking about this magical new Siri. So, a lot of people who have not engaged with AI are about to in a very real way.
And the the response to that's going to be very very cool. And so, just like when Deep Seek came out, I'm sure a lot of y'all got questions about what's this Deep Seek thing? There'll be another one in September, get ready for it. >> Yeah, I think like this might be actually one of the kinds of crunch almost like unlock stuff like I think combination of like basically open models getting good enough as well as like the on-device compute getting strong enough to serve the current of like today's frontier models, right?
Like in a year or two. Like Like basically if you can like run all posts like at a decent speeds on your like phone or laptop, I think the majority of humanity will probably like run local models. Like and and I think this probably applies more to the consumer than to the heaviest like enterprise agents, but I think it seems pretty likely to me that like there will be some flexion point even then like almost like similar to a new platform shift where like you can almost like tap into the local compute of a phone or laptop and then like start a next generation of almost like AI-enabled applications without like that that can ultimately really like leverage the local compute of like device on-device compute. >> You can run a 4 billion parameter model on your on your phone right now that is way more useful than GPT-4 was when it came out.
And I think it's important for us to continue to to to focus on how do we best utilize that in the most meaningful way possible as well as chase the newer capabilities that will come from things like, you know, drug discovery and and and scientific exploration. It's It's going to be a a fun couple years. >> Absolutely. If I were to try and summarize, I think that, you know, we're going to learn a lot more about how to these local models are are in built incredibly in the systems and their relationships as they, you know, interact with frontier models over the next panels, but this panel really shows that I think that we are at an inflection point to where if you think about how, you know, not even a short 6 years ago it was AI was really for the research crowd and not really many people cared about it.
And then of course it came into the public consciousness with ChatGPT, um but now there's this next thing which is that open source is now really starting to enter the public consciousness, but very few people have touched and played with it and have had that aha moment. And it sounds like we have the potential to do that this year and sort of guide the the future wisely, but ultimately it's up to a lot of the builders in this room as well to to to leverage that and represent, um, you know, this important inflection point that we're in, uh, on the side that hopefully brings, uh, you know, intelligence, more intelligence to all of us, which is ultimately, I think, what everyone in this room would agree is is sort of the direction of progress. >> And you know, everyone has said it on a panel previously, so I'll just also say it, which is that, uh, and and both of you have already said it, in fact.
Uh, like you you guys are extraordinarily important to this goal. Uh, every one of you who is in this room and your friends and whoever what whatever communities you're part of, uh, without you guys, uh, we we will we lose the fight, right? So, thank you for showing up and, uh, I I can't wait to see what we all build together. >> And with that, thanks Vincent and Lucas, RC and Prime Intellect, and of course Chris from Nvidia. >> [applause] >> Always, always a pleasure. >> [music] [music]
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.