Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
Words
14,291
Runtime
1:13:35
Speaking pace
194wpm
Reading time
60min
194 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
labor to the economy because we can add capital to the workers that are in the economy or because we become better at turning this labor and capital into economic output and that's what we would call productivity increases in economics and in the long run the main driver of economic growth are these productivity advances and those in turn are basically improvements in technology and also in organizational progress which is closely related and this is where I want to start and to bring in actual history um to to to draw parallel to the AI
97 words, the words spoken in the first 30 seconds at 194 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 777 |
| Average words per sentence | 18.4 |
| Longest sentence | 150 words |
| Questions asked | 53 |
| Sentences containing a number | 48 |
Most used terms
Filler phrases
727 in total: uh 218 · like 212 · um 124 · actually 62 · kind of 39 · basically 37 · you know 17 · right? 11 · literally 3 · I mean 2 · sort of 2.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
labor to the economy because we can add capital to the workers that are in the economy or because we become better at turning this labor and capital into economic output and that's what we would call productivity increases in economics and in the long run the main driver of economic growth are these productivity advances and those in turn are basically improvements in technology and also in organizational progress which is closely related and this is where I want to start and to bring in actual history um to to to draw parallel to the AI productivity puzzle that I was alluding to at the start.
We can, for example, I think very well compare AI to the prime example of a technological revolution in history, which is the industrial revolution. Um, the industrial revolution, as I mentioned, kickstarted modern economic growth. Yet, we observe a somewhat similar and interesting pattern for it too. The steam engine, its main underlying technology, contributed close to nothing to productivity gains for many, many decades after it was invented.
And it took actually over a century until we see productivity and economic growth showing up in the data after the invention happened. All right, a brief slide on technologies. In economic history, there's generally a class of technologies that's described as general purpose technologies um which are in economic history uh funnily enough referred to as GPTs. Um and nice coincidence for this talk, but general purpose technologies generally have three attributes.
They are pervasive. That means they apply across basically all sectors of the economy. The technology can be improved over a long period of time. And the third point is pretty interesting because it means that they have complimentarities. And that means the technology raises the returns in its applications as they improve and as the applications improve that in turn raises the value of the initial technology that helped bring them to life.
In other words, uh, general purpose technologies can keep on giving for a really long time for the economy. And an important question that history or future will tell is whether our AI engineering GPT turns out to be a GPT that can achieve this. Well, there's no official list. Um, technologies that are often considered to be general purpose technologies are, for example, the waterhe, steam engine, the railways, electricity, or the computer.
Um, and there's a second idea that I just want to briefly mention and keep in the back of your mind, and that's that there's other kinds of technologies that are invented that help us improve how we do research and development going forward. So these innovations can basically help us drive productivity growth going forward. And among these are for example innovations such as the printing press, the microscope, the invention of the corporate or the university research lab which we also obviously see in AI a lot right now or machine learning technology.
I think AI has a good claim to be both of these kinds of technologies at the same time. um both a pervasive technology that is unlikely unlikely to saturate quickly and can add economic growth for a long time but also a technology that can be used to improve the process of invention itself and I think over recent weeks especially we've seen that in in mathematics for example and I think that's very exciting for what it can do for living standards because as you can see on this graph on the slide um for the past 150 years is the US economy has basically grown at a stable 2% per year and per capita terms times you see the trend line fits very well.
There's periods where growth is faster. There's periods where growth is slower, booms and busts, etc. But on average over 150 years, pretty remarkably stable two years uh 2%. This is a lo axis is a log. So um every linear change basically means a doubling. 2% a year growth um doesn't sound like a lot. If you were going to invest your money and get 2% a year, you'd probably not be terribly excited. But it does a lot when it compounds over a long time.
So over the 150 years on this graph or 140 years on this graph, we roughly see income per person in adjusted for inflation, etc. grow by a factor of 10. And to me, it's pretty astounding what it takes to achieve 2% of growth. This was an era 1880 to 2020 of amazing innovation. Basically, all modern technologies that we marvel at today were were invented then. the automobile, aviation, space travel, biotechnologies, the internet, just to name a few.
And I think you can fill in the gaps. And in the end, all of this comes together to giving us an average growth in the economy in the US economy of 2% per year. Now, AI might do two things. Um, firstly, it might support this trend so that it can keep going for a longer time. And I think that's already a great achievement because we can see what level of tech technological innovation it takes to just sustain 2% per year.
But um and excitingly it might also accelerate this. There is not really consensus right now in economists um of the magnitude of what we might expect this uh graph to look like going forward. and estimates range from as little as a 0.1 percentage point increase of additional GPT above trend um per year. But there's also models of economists that take into account recursive self-improvement, some sort of intelligence explosion in AI that arrive at the possibility of high singledigit percentage point increases um to what you can see on this chart.
And as living standards compound over time, as you can see here, um what new regime of growth we end up in over the next couple of years is really going to make an incredible difference to to living standards. All right. Now, let's turn back into the history. Um and start with the industrial revolution where it all has to start when talking about technology. Um the steam engine is the canonical technological innovation of the industrial revolution.
Um in principle it helped divorce the creation of mechanical power from muscles and rivers um and made it possible to provide power at some scale and also much more location independent than we were able to do so before. Um it was first invented by Thomas Newman in 1712 and then materially improved by James Watt whose name at least everyone knows um later in the century but the technology really took well into the 19th century to continue to improve before it was widely adopted in the economy and industry and factories and only when high pressure steam engines spread and the price of power came down by many orders of magnitude we can actually see productivity and GDP take up in the statistics.
So the steam engine really had to become a radically better technology for a long long period of time because we could start to deploy it at any scale. And then it took the invention and also deployment of its compliments, new machine tools, railways, steamships, a newly skilled labor force until an entire economic system could basically develop around it. And the industrial revolution as we know it and the transformative change that we see in the numbers starts really humming and making a difference for people.
Right. I think this carries an interesting observation. Um, a technology can arrive a very very long time before it diffuses and an economic system develops around it. In the case of the steam engine first invented in 1712, productivity advances really showing in the 1830s at the earliest, so more than 100 years really until we see that show up. Another example um is the electrification of the United States. Um Thomas Edison installed the first power plant in New York Post Street station in 1812.
Um but it took about 30 40 years for um us to see any meaningful electrification of factories in the United States. And this time it wasn't the technology that needed to improve but it was really a redesign of the factory that was required as the place where power is consumed and not created. um factories had previously run on steam engines and were laid out to use one power source and you can kind of see that on the um picture on the bottom right that I made with Grockbot actually.
Um um and there was one power source that basically ran all machines in the factory. So one power source connected with belts and shafts to all machines in the factory. Um which basically meant um that you had to lay out the factory in a way to accord with how you transmit power in it. And also factories were often laid out vertically over multiple stories. Um, while you could really quickly actually replace the steam engine with an electrical motor, um, this wasn't really the best way to set up a factory that had electricity available to it.
Um, and electricity really unfolded its biggest gains when you could use its unique advantage. And the unique advantage of electricity is that you can generate it at a different point than the place where you consume it. Um, and this allowed for the so-called unit drive system to take hold, which basically meant every machine in a factory has uh one power source. And I think the advantages become clear very quickly because now if you you can turn off just one machine uh in the group drive system.
If you turn off the steam engine, the entire factory starts running. If you turn off one motor, electrical motor for a machine, only this one um machine needs to needs to be turned off. Um so you can start laying out the factory according to your workflow and not according to where the belts and shafts need to go. Um you can expand it easily. You can retool it. You can start producing something different. Um but this is a slow and expensive process. requires a retooling of the entire factory and not consumed in factories.
And after 50 years, 80% of power consumed in factories came from electricity. We'll move a bit closer to our own time right now um and starting with a famous quote from the economist Robert Solo um where he wrote that you can see the computer age everywhere but in the productivity statistics and he said that in 1987 where personal computers were already all over the place and computing quite widely adopted in the economy.
Um and I think in general very well known uh second half of the 20th century we saw really amazing innovation in computing personal computing internet I think altogether information technology that drove down the cost of processing storing communicating communicating information precipitously but it took until the mid to late 1990s for productivity to actually respond here and again the economy needed some changes and this time these were mostly organizational so firms and workers need to change their processes um develop the skills, adopt the right compliments to really gain from computer technology.
And this time around that was things like developing well other firms had to develop enterprise software, database technology. Supply chains needed to be redesigned to actually work with digital ways of accounting for for for goods. Um and of course also it meant the adoption and rollout of many many many individual pieces of hardware and software to workers in the economy. And this is actually a much slower transformation of the organization that was required rather than the invention just of the computer technology that ultimately helped us take advantage of this technology in a really meaningful way for the economy.
And um this relationship or this kind of development that we saw there is sometimes described as a J-Curve relationship where um the investment required initially to adopt the technology to start learning new processes, buy new hardware etc etc actually slows down productivity growth before it enables it to to accelerate um later in the process. And I want to very briefly touch on another place where technology can be out of sync with another factor and that's in capital markets.
Um, in history, the real promise of groundbreaking technologies has always been able to marshall significant money and resources early on before we actually know how and when a technology is going to produce growth in the economy. And I want to put two historical episodes next to AI's buildout and compare it a little bit at least in magnitude. Um in the 1840s in Britain, uh Parliament authorized hundreds of new railway lines um that they could be built.
And this was basically happened all at once. And a sort of mad rush of cheap capital, um rising demand for freight, especially coal, and a belief that every town and everywhere in the country needs to have a railway connection to to participate in the economy led to a crazy investment boom that has not many many comparables in history. um it's one of the largest of all time and railway investment reached 5 to 7% of GDP um in a single year in the 1840s.
And another example uh closer to digital technologies from the late 1990s um the computer and the internet in the late 1990s commercial internet takes off and a race to deploy longhaul fiber started to basically enable people broadband access. Um and there's a lot of deregulation in the US. New carriers, new companies that were laying fiber and expectations that demand and transmission of data was going to to grow insatiably led to basically a way overbuilt out of of fiber technology.
Um this boom was around 1 to 1.5% of GDP in 1999. And today if we compare these magnitudes to AI, um AI infrastructure investment is already at about 1.8% of US GDP this year. are already bigger than fiber and forecasted to reach about 3% of US GDP by 2028. So in other words, uh AI infrastructure has already up there with um some of the largest investment booms that we've seen in history of all time. I'm not going to really be drawn out to to comment on whether it's a bubble or not.
I really don't know. uh but I think there's an interesting observation from history that no matter what the the private returns for investors end up being from investment in this infrastructure the infrastructure that's built out can outlast and turn out to be productive later on even if there's a bubble. So this is the case with um railways surplus railway tracks that were late and overbuild or it's also called dark fiber um in the com era that later enabled much cheaper transportation and things like streaming the mobile era to to happen with much much cheaper infrastructure available to the economy.
All right. Um a strong pattern that I've hopefully driven home um is that technology can move at two clocks. Um, one is the faster clock of technological innovation and prototypes when a transformative technology emerges and the second is the slower clock of actually deploying it when a technology gets deeply integrated embroidered into the fabric of the firm and and the economy as a whole. For steam, that was additional technological innovation in the core technology.
And for electricity, it was a new factory that needed to be designed to complement um the new technology. And for computers, it was basically large scale organizational organizational change, new skills, new ways of working and also the deployment of course of hardware and software. We're now four years into the generative AI era, if you will, and I would like to invite you to think about what it'll take for AI to ultimately start to raise productivity and economic growth for everyone.
Um, is that the core technology improvements in foundation models? Is it adoption across other sectors of the economy that are currently not really touched? Is it complements the application layer? Is it reskilling the workforce? I don't really have the answer with me, but I'm excited to speak with everyone uh that wants to approach me over the next two days about this. All right, to wrap up, um I hope I've made the point that invention creates possibility as a start.
But to create turn this possibility into growth, more is needed. diffusion, investment and compliments, organizational change can all determine how much of this really turns out into productivity which turns into improved living standards for everyone. Um, bringing AI into firms, rewiring how work gets done, training and transforming how other people work, and starting new companies that do this from scratch with new ideas is, I think, very much how this technology gets on the road.
And that feels like a very encouraging message to to end on for this room because I think most of us spend all of their days working into rewiring the economy for this new technology. Thank you so much for your time and have a great conference. [applause] Let's give it up one more time for Clemens, please. Wow. I always like to see how we can, you know, extrapolate what happens in the past to kind of try to predict the future.
And uh Clemens spoke a little bit about how we should feel about today. And our next speaker is going to talk to us about the future. Actually, our next two speakers are going to talk to us about the future. And uh our next speaker is somebody who doesn't need any introduction. If you've been on social media, if you've been on Axe, you probably have seen his hot takes. Uh he calls himself a master tinker. I still don't know what that is, but we'll go with that.
Okay. All right. So, please let's give it up for the pizza. Okay, this never works, but let's try to make it work. All right, slide. Boom. You guys can see it. Perfect. I just realized how bad it sounds. We call it master tinker. I just like tinkering with stuff and we form this entire thing that I'm going to tell you in a second. Are we going to stop the music or am I going to have a background soundtrack? Okay, it's going away.
Cool. So, I'm the kids on X. If you know me on X, we probably disagree about many things and we probably argued about many things. So, um I have 20 something minutes before I need to catch an Uber. That's going to be fun. Um I just need to get my cursor over here so I can click and it will be perfect to start. So, usually I plug in my 39 products I'm working on. This time I won't. I'll just have one slide. If you are someone who tinkers with like local LMS and hardware and self-hosting and a bunch of other things, we have a group of crazy people called thinker.club.
That's the URL. And basically every person within this club is a copy and paste version of themselves. It's hilarious to see that everyone has the same interest from like hardware and 3D printing and AI and SAS and business owners. If you want to join us, it's at thinker.club. So previously at AI engineer, I've given a talk uh called from vibe coding to vibe engineering which still many people are going to that transition.
So if I would do the talk right now, it would still be relevant because so many people are just moving from vibe coding to vibe engineering. But I can already sense the shift within me within my own apps and within other people's like I can see what's everything that's going on on Twitter and we're slowly starting to move to the next level of engineer nearing you probably heard the term like factories and software factories and slopp factories that I call them and you can see that within Q4 of this year and within next year it's going to be more and more and more about going from like VIP coding and even this old version of VIP engineering to something new.
But before we talk about the something new we got to go a little bit back. So I'll take you just a few slides through what we went through in the last few years. There's still many people who are using copilot and they're pressing tab and they call that AI. But we went through that for the longest time. Uh at some point we were copy and pasting from chat GPT. And if you remember if you're old enough like few years ago it couldn't finish a function and you have to beg it and be like can you please finish that and you wait for the next one and you copy and paste it somewhere.
Then we had to beg again but that time we were begging for JSON. We're like, can you just not put markdown in it? Just I need clean JSON. And it be like, sure, here's some markdown and some JSON. So, whoever started with AI now, you're very very spoiled. Then we uh we focused on editing one thing. So, there was no agents, but we would just use AI and it would change a file and we were like, oh, this is dangerous. It might do something, right?
And then very quickly, it felt like we went through agents, but we were still careful and we were just reviewing every permission. We're like, be careful. Don't touch the the thing. And then like literally one year later, like when it asks me if I can run this, like how the should I know, right? You just read the thing, you're like, "Yeah, I I guess I I hope so." Right? So now everyone is in this full yellow mode. There's not a single person that I know that's not actually running on full permissions.
Like take my production databases, customer data, GDPR, whatever, man. Take everything. Just finish the the job. So we ended up here. Basically, we have a few categories of people that are working with AI right now. We have the slop grenade throwers, which is most of us. We're trying not to be, but we're just throwing the slop grenades. We have the meat proxies. Have you heard this term? Like, this is one of my favorite terms I've heard lately.
Like, I absolutely love it. And we have actual AI engineers, which are like not a lot of people, but we're trying to make like uh everyone an AI engineer basically. the slop grenade thrower. It's like everyone is kind of doing this that because of the speed of AI, it's very easy to just prompt an agent and then it generates a bunch of slop and then you don't read it and you just throw it at your colleague. You're like, "Good luck.
Good luck. Hopefully it doesn't explode production. You figure it out." And then we have like the meat proxy which is like if you don't know what this is it means you're sitting between your manager or your supervisor whoever it is and they ask you to do something and then you literally take their prompt you put it into an LLM then you take the result from that LLM you barely review it you give it back to the person so you're just serving as a meat proxy and currently there's so many companies and this would have to change that they're actually consisted of layers or layers or layers of just meat proxies talking to each other and I think next year is going to be brutal because many people they're going to wake up to the fact they're going to go to the mirror and be Oh And they're going to get fired.
Like it's inevitable. I've been talking in my talks about this and it's going to happen I think around next year. The actual engineers actually try to build systems around their processes and no one has the right system. So this talk is about like my road there. I can see like two camps of people. It's always been a division. So we have the desktop people and we have the CLI people, the hackers, the tinkers that that like I always hated the terminal by the way.
I always hated CLI. I was a big hater. Like give me a beautiful UI that I can click on. I always hated this tinkering type of stuff. What's mindblowing to me that we had like three years of orchestrators being released from cursor to conductor to everything else. Everything like no one is even trying. It's just a copy and paste of the same layout. On the left you have sidebar with chats. In the middle you have chats and on the right side you might get a git diff and everyone is like we just raised 37 million for a new orchestrator and the screenshot you're like is that codeex?
Is it cursor? is it's the same always and no one is even trying. I'm going to try to show you mine eventually at the end of the talk. So, we'll see if you like that. Uh CLI camper the vanilla once and I'm going between back and forth like I launch a desktop app and then I go back to the CLI and then we have like the absolute crazies who like if you haven't heard of the term oh my pi don't even research what is oh my pi but we have people who have these oh my pi inside of her inside of inside of VM inside of proxmog and ask them what they shipped and they be like yeah but my factory works I guess.
Um, so I was the one roasting all of this. I always hated like self-hosters. Like I was always the one just pay Versel, pay railway, pay whoever, don't bother with this. And then I just full went and do the other side with Cur Club because the more people post about this, the more I want to do it too. So I have like Linux boxes and tails scale and a bunch of basically. So we have tools versus methods. We're going to talk about methods.
So So we have like the two tools like CLIs and desktops. And then we have uh methods. So the one me like I've been going through all of these and then I'm going to tell you which one that I landed on. one chat and prompt until work. This is like I would stop opening multiple chats for different tasks and I just keep one and especially with the smart compaction of codecs. I can just dump a bunch of things in the queue and I don't even know where it compacts.
I think it doesn't even tell you that it's compacting. So it would just work and pick up the next task and pick up the next task and the next one and it would work and ship but in between every task there's no like management where it's like launches a sub agent to review and then another one to test. you just hope that it's going to go through all the task and it's going to finish all the tasks. Kind of like it, kind of dislike it.
Then we have this is actually uh one chat per task where you have like a folder in codeex or whatever and then for every task that you need to do, you start a new chat. But if you don't use work trees, you either have to guess whether the features are going to overlap each other or you have to use work trees and then you need a chat on top of those chats to do something with the work trees. So it's kind of getting messy.
Then you have like work trees and hope. So no, nobody knew about work trees by the way until like seven weeks ago. So shut up. So now everyone's using work trees and everyone is a git engineer and you know stuff, right? But the problem is these are still not fully isolated systems. They're like they're a different git thingy but they can still nuke your database equally as like not using a git work tree. Then we have all of these terms which a lot of times I hate because they're overwhelming to learn all of this.
Like they have, have you heard of gauntlet loops? First we had loops and everyone was like what? And now it's like, oh, I just did this using a gauntlet loop. Like I keep we keep inventing all of these terms and some of them are useful, some of them are not useful. So yeah, and finally, we're talking about software factories. If you ask any of your bots to make a research of all the articles published about software factories, your mind's going to be blown and you're going to read through half of it and be like, "Oh, just let me prompt for sake.
Let me be a meat proxy." because it's very overwhelming because uh if you look at if you zoom out like almost nobody has an actual software factory and people are just scraping it together and we don't agree on what a software factory should be. Should it be on a level of a company organization? Should I have as an engineer uh um a software factory and we end up with this slob grenade fight. So if you look at any GitHub repo, any organization, you just have someone who opened a PR using an LLM.
You have a beautiful readme with like emojis and paragraphs. We never even bothered to make something bold on GitHub, but now we have like paragraphs and dividers and everything. And then their colleague responds with even more beautiful review. It's even longer. They're paying attention. Suddenly things you're like, you didn't write that you You know it's by an LLM. So now we're going between like two meat proxies serving two LLMs and they're talking to each other and nobody realize and eventually the task has been submitted either by a customer or by a superior, right?
And no one is realizing what what what are we doing here? We need to like untangle this. It's actually predictable in some way to know what are we we doing. So I throw one back. My AI reviews your AI and blah blah blah. Eventually they meet in the middle. But we need to find a solution and we need to make the process to evolve. So my road to this from terminal to desktop apps to terminal again. I kind of recently went back to the uh CLI.
Who has heard of this? Heard of this? Get it? Huh? Heard of it? You get it? Okay. Okay, so if you don't know what this is, it's basically like a T-Max thingy that if you do remote coding on any of your devices, the sessions are not going to die. And in the terminal, they're going to be very nicely organized so you don't lose track of of what you're doing. And um ironically, like I've been going between so many setups.
I even made my own orchestrators and stuff. And like last week or two weeks ago, I just set up Herder with like the simplest terminal and I shipped the most. And I'm like maybe we were over complicating the entire software factory thing. But it really depends. Are you working with people? How many people? How many teams or are you solo or shipping? Shipping what exactly? I moved to Pi. Pi rocks. I like Pi. I was making also fun of everyone who used Pi.
I was like just be a person. Use Claude, use Codex, use whatever. And then as my quotas were going, you know, like my claude quota would die. And then I would switch to Codex CLI or Codex Descope, whatever. They have completely different behaviors like how they launch sub agents, how they do compaction, how they do this, how they do that. So I would switch between the cursor CLI, between cloud CLI, between codec. And as I switch between them, they completely have different behaviors.
And this was driving me crazy. So when I build an orchestrator, I build it on top of the three, but it's so hard to build an orchestrator that needs to serve like it's not only the CLI, it's like they have different methods of doing things. And with PI, you just have a CLI that does nothing out of the box. So if you wanted to compact, you need to install a plug-in. If you wanted to launch sub agents, you need to install a plug-in.
And you fully control the process. So it's like a Linux of harnesses, which I see none of you are like, "Oh yeah, yeah. I would love to do more Linux of things, right? So, Herder also has this nice feature uh I think they added it recently that if you have Herder on all of your machines, when you open it on any of the machines, you're going to see your I don't know MacBook Pro, Mac Studio, Linux, VPS, whatever, and it's all in one place.
So, it's kind of like very nice. I shifted this year to using my laptop only as a remote control. So, I started by the cloud and whatever and now I'm moving to my own things. But, it's this laptop or my phone are only remote controls for things that are happening in the cloud. So I went into this deep rabbit hole of like I don't want my code because I spilled milkshake on this laptop and I had like an unpushed GitHub repo and thankfully they saved it but like I thought it's going to die and like I don't know like a month of work would have died with this laptop and I'm like I need to move all of my work somewhere and just use any laptop.
I even used my 10-year-old MacBook Air to just remote into my setup and do things remotely. So I started first by shoving things into the Mac studio but I realized that Mac and Mac OS is not made for remote agents. So then like probably after 15 years of not even thinking about this I was like fine I'll just try and maybe Linux will be a good thing. So I started renting a VPS for like $128 per month and I was like ah if it makes me money it doesn't matter whatever.
Then I did the calculation. I'm like but if I pay this for two years I can just buy a Linux machine and suddenly you have a Linux in your basement. But I realized I have a gaming PC and I don't like Windows and I don't like gaming because this is way more fun than gaming nowadays. I don't want to game. I don't want to watch shows. I just want to tinker with the with the agents. So, I just decided I'm going to format this PC.
I'm going to put disgusting Linux on it, right? But Windows is like up there. And then I realized that this works. It's great, but I don't have enough RAM. So, I started searching for part. Like, I went so far away from actually shipping value to my customers. And the rabbit hole goes even deeper. So, I was like, this is 32. I need to get more, but I need to upgrade the motherboard, but I need this, but I need that. I ended up with a second Linux PC.
So, I built a second Linux PC from scratch. And then someone was like, why do you choose which Linux to DRO to install? You can install Proxmox. Now, who is nerdy enough here to have set up cuz I'm going to skip this. Okay, we have a bunch of hands. So, this is really cool thing that I used to look at all of these things we discussed because I had to I had to go into the configs and click thingies. Now, anything that I mention here just assume I'm not smart enough to deal with this.
I just told my agents, hey, my friend told me to use Proxmox, use it. And then boom, it's it's getting set up on the thing. So this is nice because actually on one computer like let's say you have a PC you can have this is like a higher level OS you can have multiple Linuxes you can have a windows you can have Windows XP all of them as like virtual machines and then your agents do everything beautifully you don't even need to manually tinker with anything they also have these things like VMs and containers so basically you can like self-host on your own computer a bunch of things that are available on your network and you can actually run full VMs so my agents like my Hermes agent and whatever I've given them small VMs so they cannot break out.
They cannot do something stupid. They just have like mini computers within my uh computer. So, how I promified, I just plugged in a KVM and I went to Codex and I was like, "Please finish this. I don't want to look at it." And then it worked, it worked, it worked and I saw it and I was like, "Cool, it works and I don't want to um touch it." So, a lot of times people start remote coding and I see like people who send me screenshots, they have some weird IP addresses and So, you can do like a very very nice setup.
If you're remote coding, you can just tell your agent to set up traffic, however you pronounce that, rayic and PM2. And then all of your processes are going to be like magically managed and on nice URL. So you can have your blog, whatever your domain is, and as subdomains, you can have like dev processes. So anytime you want to see how your landing page looks like in dev, you can go to my beautiful landing page-dev.m my URL whatever and you're going to see it and you're going to remember it and you're going to bookmark it.
It's not going to be a random IP address. So that's like a nice uh tip there. Agent sets it all up. So the entire point of this, people ask me, "What's in your config?" And I'm like, "I don't know. I haven't seen a config in two years. I just wish for things and the things magically happen." Then I hit the limits because of course all of us are dealing with this. I have a tip here. I'm paying for four $200 subscriptions.
It's still cheaper to than a human. It's like kind of fine, but you need to manage them in a different time. So there's this thingy called codeex load balancer which you can like what you got you guys don't use a load balancer. [laughter] I thought this is going to be like widely accepted and people going to be like yeah yeah I do that. So this is basically like one endpoint but when you hit it it chooses from which one to siphon from.
So it knows your limits, your weekly limits, your daily limits, whatever. And it just like my limits never expire basically until they released Astra and then they were gone in a day. So with Astra it was like a At bonus, you can do live voice on the same endpoint and you can generate images on the same endpoint. So you can give this endpoint to all of your apps and they just use the load balancer and it gets worse guys.
Um then you have like that's for codex but what about the other ones? I have like cursor, gemini and a bunch of other things hidden behind this easy CLI proxy API which is like another load balancer proxy. And then on top we have an AI router which is like the master endpoint. So when you when I want to code I talk to it and then it figures out who to ping from all of these things and I still haven't shipped much. So as I'm going through the slides just be like you can be more productive with just like pi herder and and whatever.
So every model one endpoint everywhere it doesn't uh create new capacity just trying to squeeze out from a little bit of claw a little bit of codeex this and that to actually make it work. So my connections were scattered and one problem here and like this would be one slide that I take a picture of just so you remember the name exeutor if you haven't set up exeutor. I've been spreading the word around more than I did for openclaw because this is the like it's absolutely the greatest because I would want to switch setups.
I would want to try I don't know Grogbot or the Madeam Muse or this and that and anytime I would have to connect my Gmail, connect my calendar, connect my GitHub, connect my Cloudflare, connect like all of the connectors and what this is is like a house of connectors like with one MCP. So I've plugged in everything to my exeutor and now if I want to try Muse by by Ma I can just tell it to use one MCP that's my exeutor and through the exeutor it's going to ping all of my other services.
So it makes it super easy to switch between I don't know pi claude codeex any of the bots. It's really nice. Then I realized uh this is open source. It's on my GitHub. Um it's called skillbox. I realized that we're doing it wrong that our skills are stored in markdown files. Like I had like a GitHub script that sync my skills across spaces and it was very messy. And I would like set up a new computer. I'm like I don't have my skills here.
And I'm like if we have exeutor for MCP and exeutor actually holds all my MCP, why I don't have a single MCP for all of my skills? And then it turns out as I was working on this, they made it an official spec. So now in the MCP spec, you actually have something called skills over MCP. So instead of having 30,000 markdown files here and there, you can just self-host your own like you can tell your bot, hey, I like this skill box.
Set it up for me. It's going to give you a nice URL with like a graphic interface and everything and all of your skills are going to be there. And not only that, there's a thing called skill bundles. So you have like a personal agent, a coding agent, and other things. My personal agent needs like four skills here, and my coding agent might need another 20 skills, but they don't need to know about each other. So if I set up my open claw or Hermes, I'm going to point to my personal skill bundle and it's only going to work with those and it's not going to be like, oh, what is Cloudflare?
So that was very annoying that all of my skills were for everyone and in one folder with a bunch of markdown. So now my setup is portable. I can just take it and just plug into something else and it works. I'm currently working on something new which I'm not sure if I'm going to finish. It's like a layer on top because I realized like we have skills. We have MCP and Codeex. The Codeex app calls the layer on top. They call it plugins, right?
Right. So if you go to Codeex and you install like a Mac OS development plugin, it's going to install MCPS for you and it's going to install skills for you. So I'm like we need some new names like I don't know like playbooks and like I have a bunch of ideas. Maybe I'm going to present that in the next presentation. You're going to see it on my Twitter. But the idea is like I can see skillbox and exeutor kind of merging and a couple of layers on top would make it way nicer to pack your entire setup and just plug it somewhere else.
Like Kodi by Keny Dods if you've seen Kodi is like kind of that. It's hosted on Cloudflare workers and a bunch of complex terms that I don't actually understand, but his idea was basically to take his brain of his agents, put it into Kodi, and then he can just port it anywhere he wants. So Cody's like a skillbox and exeutor with a bunch of goodies on on top. And executive is coming out with a new version, version two, which is going to be similar to Kodi.
It's very hard to keep track of all of this, but I hope I'm going to plant some seeds that going to tell you where things are going with this. So our RAM, our our brain actually cannot keep up with everything that we're doing. So we have power, we have the four subscriptions, we have a load balancer, we have a proxy, we have this and that and I decided that unless I build like my own orchestrator to organize all of this mess, this will never go anywhere.
It's going to be very messy. So my from everything that I mentioned to you like how I was actually working through things, this is what I landed on right now on this day today. So I landed on realizing that one chat where it constantly compacts and I dump a bunch of tasks is not going anywhere because there's going to be plenty of conflicts. Then we have the other thing where I have multiple chats, they're going to trample over each other.
Then we have the third one with work trees. That's also problematic. So I realized there was this article, I hope it's in my slides here, that you get one chat, then one work tree and one PR. So this is what many people were doing. I personally as a solo indie hacker let's say I wasn't even working with PRs I would just ship everything on main and it would go in a cycle but now I realized I want to set it as a goal to have like a very high quality PR with all the type checks with linting and a bunch of things and I want to basically automate the process there instead of doing things manually because it happened like I started using only work trees and one of my work trees because it's just a g thingy it wiped the production database luckily I had a a backup and it's like oop sorry I was trying to see the database to test but I'm not testing in a container I'm just testing on your machine and it just basically wiped everything because a work tree is not an actual sandbox.
I have the slide here orbs. So this guy from AMP code thren something he was talking about orbs but my brain was so overloaded by these new phrases that I'm like you. Like orbs are not happening and I just tweeted that a couple of times like orbs are not a thing. They're not going to be in our dictionary. Shut up. Don't talk about orbs. And basically his philosophy was that anytime your agent works on a task, you need to give it like a mini sandbox, mini VM, mini computer thingy where it can work in it.
And as it works, it exposes a tunnel with like one URL. You can observe that URL, you can scroll it, you can check it. Basically, you can be the tester, right? Our role is a bit changing right now. We will we're going to be the testers for the AI. So, at first I was making fun of it. I was like, "This is not going to happen." And then I went through the entire journey that I told you and I'm like, "Orbs, it's orbin time." He had that phrase and it was like cringy, but it is orbin time.
And I realized that this is the most stable way to actually to make something that's properly tested. that's isolated and that's been gone through like a software factory process. So I have like one task making one throwaway mini machine. Remember how I mentioned Proxmox? So my Proxmox device can spawn like it has like a little docker thing. I cannot believe the words that I'm saying. Like if you watch my talks for three years ago, you would be like who the is this guy?
So it makes anytime it needs like an orb, it spawns a mini VM and that VM is like has everything. It has the database, the seeds, the browser, everything that it needs and then it just works um in that. Uh, I'm using something called Podman, which honestly I haven't heard of until last week. And it cannot take things from outside like uh production credentials and blah blah blah. And it has like a queue, so it doesn't destroy my computer because it can be like only eight orbs at a time before my computer dies.
So they just queue up. If I need to do a new task, they wait and one orbs comes out, it dies, and then another one gets spawned and it they go in a nice like cycle. What's nice about this, and I've mentioned this in my previous talk, that when I was working multiple tasks per chat, there was no not many guard rails. And even when they were, you would go in your ESLint and you see half of it is disabled because the agent just wanted to get done.
It was like done, no errors, are you? Yeah, I just disabled five of your configuration lines. So, unless you put these very hard guardrails, the PRs are going to be slop and you're going to be a slop grenade uh thrower. So, I have this combo. This anti-slop is a library by Dylan um Monroy, Mulroy, I don't know how it's called. So, we have Oxlint, Ultraite, and Anti-Slop. These three work together. You can also take a picture of this, tell your agent to set it up, and they make sure that every PR is according to these rules.
Because you have the rules in your head, you need to get them out of your head, put them in a llinter, put them in a thing that checks them to make sure that actually you're going to get the result that you want. I realized that many of my rules don't fit in oxintable. So, I just asked my agent, can we just make our own police CLI? Let's call it police. And if I have a specific thing like never use your own buttons but always use chats and buttons that you cannot put that in lint but my police thingy like my agent was writing things and now it has like many many many rules that actually go on top of the linting and they make sure the agents don't do something stupid and now finally we have Jeff and Jev is pretty cool and I realize why I made like Jeff rabbit.
It's on my GitHub and it's like a code rabbit, but you describe plain English rules of things you don't want to happen. And Jeff very quickly and very cheaply goes to your entire PR and checks if your comments are useless or you're doing something stupid like instead of writing LinkedIn rules, you can literally express yourself in the English language and for every rule that you have, Jev is going to go and check that PR.
So this is like a one extra uh level on on top. So I built my own factory which I wanted I really wanted to show you this how does it work. It works like with all of these things. Everything that I mentioned here from the load balancer to spawning orbs to like testing and everything is part of the thing. And I'm looking at the timer. Unfortunately, I won't have time to show you. Well, maybe I'll have like a second or a minute to show you how it looks because like I want to just tell you that um these factories can actually look different.
So, just give me a minute here. I want to switch to this will be the first beautiful agent orchestrator that you see in your life. If it loads, it's going to load. So this is not like anything you've seen. Well, maybe on a wide screen one. It's like if Steve Jobs woke up and he was like, "Yeah, we're going to make an orchestrator, right?" So it's not a left sidebar. It's not a bunch of like things that you've seen. It actually organizes all my projects in a nice way.
So I can go, let's say I'm going to go to my bench project. I'm going to go in here and here I have a bunch of to-dos that I want to do. Uh I go here, there's like a plans and heartbeats and a bunch of things. And there's a factory which if I click, you're gonna see these are the processes that every task goes to. So if I add a task, we go through the soulbot which is like coding on the thing. Then we have soul fast which is reviewing the thing and they're going in a cycle until it's like fully reviewed and then when it's fully reviewed we have the tester which clicks around and checks if the thing is done.
So, if I go back from the factory to one of the projects and if I want to get something done, I just go in the to-do here and I can say, I don't know, the do the thingy and if I press enter, this won't work right now, but it should jump into ready and it should go through the entire factory process until it's in review and then you get like a beautiful interface where you review it and you actually feel like you're coordinating a factory and not just talking to a bloated orchestrator app.
I am out of time. I wanted to show you this more. It's called Bench. It's going to be announced on my Twitter hopefully. bully me please soon because it really changed the way I don't use CLIs I don't use codeex apps I don't use anything I can open it on my phone I can open it anywhere and just from any underpowered device I can orchestrate my entire slop factory of things so thank you for listening I'm out of time I have a flight to catch and appreciate you listening to my rambles thank you [applause] thank you let's give it up thank you all right let's give it up one more time for the kids Wow.
[applause] I forgot to say that the kit is one of uh actually the most viewed speaker, the AI engineer. Like if you go on AI engineer on YouTube, you'll see so many of his talks and they're like they rank always among the the best and the most viewed. So um it was also interesting to see how u I relate to many of those problems in my day-to-day. I forgot to say that I work for Replet and one of the things that we do is we we we we build you know infrastructure to to to uh to to uh build applications as well.
Sorry, I'm going to get that right at some point. All right. So, we are going to still talk about the future and now we're going to hear from VP of engineering of Mistral and he's also uh the third employee of the company. So, please join me in welcoming to the stage Leo. All right. Hello everyone. Thank you so much for uh coming here tonight for being so many. Uh today we are going to talk a bit about what Mistrol has been cooking uh what's new uh and hopefully give you a few information about what we are uh actually been working on.
So for the ones who do not know us uh we are Mistrol. We are the top frontier AI lab in Europe. Uh we started three years and four months ago. Uh very small company at the time. It's now 1,300 of us. So we grew a lot. Uh we have a clear mission putting AI in the hands of customers, changing the way enterprises operate. So we're not looking for AGI, we're looking for usefulness and this is why we now have a team that is dedicated in making things uh on the ground turning PC's into production.
And this is uh basically what the company is doing now. So if we look back at the previous AI engineer summit a year ago, I was um very excited to be invited last year by COB uh some things happened uh since then. You are more than 700 to be registered. We have more than 50 speakers uh today. So it's an amazing event. And about COB um that was inviting us uh last year uh we now had the chance of having them within Mistrol.
Um, and you could ask why did we make this acquisition? Well, as we are moving towards agentic workloads, swarm of agents going everywhere, it made so much sense to get a team that focused on building this kind of agentic workloads completely different from getting servers and things like this, just focusing on the sandbox, focusing on the unit of compute that you need to actually run workloads at scale. Um, this is what Yan said uh last year.
If it works from day one, we've stayed close to the hardware to provide high performance with sustainable costs. We are operating tens of thousands of applications of on top of bare metal servers across 10 locations worldwide. And they're joining mistrol to build that infrastructure. So making sure that for our own workloads, anything that we need to build and that requires secure, safe and fast async uh execution in the background, we could leverage their technology.
We've also in May announced the acquisition of Emmy AI. Um Emmy is a company that is focused on physical AI simulation. Um, and if I read the same quote from Yannes, u, by integrating our expertise into Mutual AI's world-class AI ecosystem, we're positioned to revolutionize core R&D. Together, we're providing the foundational intelligence required to design and build the next generation of aircraft, vehicles, and semiconductors.
So this is really about bringing all the capabilities, the simulation, all the things that actually help understand and model the word back into the models so that we can actually help our industry customers be better. We announced earlier this year uh $150 million partnership with Orbus and this is typically where having this kind of expertise inhouse will help us build state-of-the-art specialized models uh for these kind of customers.
We also announced uh last week our series D. This is the biggest round from a technical company uh in Europe. So three uh billion euros. Um we are very proud that this is also gathering a lot of investors from Europe from all over the world who trust us with this mission of actually making AI useful of transforming uh companies. And this is showing that uh the way we've been working uh for the past three years is actually working.
But I I [cheering] I'll do it again. You you watch me. I didn't do it. If it happens next, you you'll know that it's me. Um I'm going to talk briefly about the future of work and the way we see it. Uh this means how we also deal with agentic safety. Uh we're going to talk about sovereignty a little bit and then how we are uh working towards new models. I know you all came to hear about lhaton fat. I'm not going to announce it today uh but uh rest assured that we're working hard to get it there.
So um this is a tough one apparently. Okay, perfect. Um, so the future of work, uh, let me talk a bit, uh, about our product suite today and how we envision it. We, uh, do have a strong product for developers called the studio in which we built a lot of primitives uh, on top of the models themselves. So you have so many models today and that's not enough to go into production. And that's why we developed the studio uh which brings a ton of product value on top of this uh things that you actually need to go into production.
So think about workflows like long lived execution, the sandboxes that I mentioned earlier uh all the primitives that you need to actually make this happen. uh and we do have vibe uh that is our end userf facing product that leverages all those capabilities uh to actually uh well solve problems for productivity for uh end user customers uh and especially enterprise users. Obviously we also have a strong solutions team that is building amazing use cases for all our customers.
Uh and we are also having science team obviously building models generic models but also specialized models for everyone. But what we really want to get to since I have no power here, sorry, is there uh so this is Steven. Uh Stephen is a very nice person. He works in finance at Mistro and he is in charge of making sure our revenue is properly used. Um Steven is what some could call a tinkerer, but he doesn't have time to build uh new workstations in his basement, unfortunately.
So he needs to have a tool that will allow him to ask very simple questions but also build applications. He needs to have access to tools that will let him uh build for instance an application to analyze and cross-check some data coming from the data lake from the Excel spreadsheet from the uh notion uh workspace and make sure that he can wrap that up into an app that he can share to uh basically his co-workers. So this is pretty much what we heard before but he operates in a corporate environment where there's some guidelines where we need to make sure that the data is properly used and this is the unified version of vibe that we're uh actually working on is how we make sure that this kind of users of builders who are just regular end users but starting to know more and more about applications about how to use those models to create useful stuff are going to leverage uh our product suite to do so.
So this is typically an automation that we've been working on uh mostly for the engineers uh how we start from a slack alert that something is wrong in the middle of the night to an agent investigation. So making sure that basically the on call engineer that is getting paid while he wakes up while he goes to his computer has ready to go information to address uh when they try to address and process the alert. And this is working typically with um our vibe stack because we'll have immediately asynchronous agents running on some boxes.
You got it. Um that will plug through the connectors that we have graphana connectors, sentry connectors, uh GitHub as well to understand what PBS could have regressed all of that and propose uh some investigation some clues that will help the engineers that is on call uh among a huge system to understand where the issue could be and then obviously post back the findings on Slack. uh so it can be a conversation and so the UN call engineers are able to start from somewhere without wasting too much time.
So this is uh an example uh you'll have the initial sentry alert uh and immediately the agent is answering with some of the things um that uh it found uh that will help the engineer understand what's going on. What's funny with this is that you used to need to do a lot of plumbing to do this because you're plugging multiple systems. You need the um agent to be running somewhere. You need to be setting up those connections saying you need access to this connector and to this other connector.
And this is why we're focusing on this um unified vibe version is where you will start from this kind of prompt that is simple enough for our dear Steven uh to enter it. When interview feedback is missing, send a personalized reminder to the interviewer who forgot to do it. and before you need to think about connecting the right MCP of your ATS or something. What we're working on is making sure that this is pretty much everything that you have to do because all we've just seen in the previous presentation is entirely right.
Um, and you need to think about it. But for our own tinkerers, enterprise tinkerers who are actually not developers and who actually know nothing about Linux VMs or something, they just want to be able to send such a prompt and have all the connectors and all the enterprise tools that they need readily available. And this is where VIP can now understand this intention, spawn agents in the background, build some code, run it, run the test, deploy them as well, host them, and basically iterate by itself for minutes or hours till you have an end result that you can actually share to your co-workers with a set of, you know, restrictive u access controls based on your permissions.
All you need to actually make enterprise software. And this is where it gets interesting because this kind of use cases you see them everywhere and everybody is code coding or you know connecting whatever they need by directly plugging connectors just letting any kind of data access any kind I mean any kind of agent access any kind of data without any kind of control which basically prevents this kind of use cases to be deployed into production.
So we really think that the future of work is making sure that we bridge um all those personas into one and this is where we're investing. So instead of using software because our dear Steven does not necessarily know how to make software we just direct agents. Um we have one unified harness. So I'm not going to present a different product if you're in finance or if you're in legal or if you're an engineer. the product, the model is able to adapt, what it will use, what it will show with a unified harness across the different surfaces.
So, what we saw earlier again is exactly what we're working on, making sure that you can start a session locally on your own laptop and then teleport it to the web and then execute it in sandboxes, just check it on your phone, on the app, actually steer it. It can actually push to GitHub. It can actually check your CI. It can iterate on that. And this is where the one unified harness works so well because a business user who just wants to process some spreadsheets will be able to use the same without noticing the difference.
So the platform obviously gets more uh valuable each time someone uses it. Someone create a skill, someone creates an app, they share them with the rest of the team. Um we can also refine the skills, refine the prompts. So typically in the use cases that I've just shared uh well you can see after one week of execution uh of this investigative agent what worked what did not maybe the way it called the graphana MCP at first was wrong because it did not understand this undocumented option well by looking at the traces now the skill can self-correct and can actually specify how to properly use this connector in the future to make sure that it will be more token efficient and faster.
This leads to agentic safety because those agents, they're so powerful. They do so many things, but obviously they have risks and I'm sure you all have someone in your family that in the past 10 days ask you if AI would kill the word. Well, Mrol will not and we are investing heavily in agentic safety for that reason. Um, briefly, what is uh agentic safety? Uh if we talk about the different threat models that we have with agents, they usually have privilege access to a lot of connectors, a lot of data, and then can do anything with it.
Since they're nondeterministic, it means that it's really hard to put hardcoded uh guard rail uh saying if you do that, well, don't. It's hard because next run it might not do it actually. Uh and then obviously attacker control inputs the prompt injections that we all know about uh they're getting more and more sophisticated and you know we try to uh make it and bake it into the model but then each model has a jailbreak and as soon as you're fetching uh user data or data coming from the outside uh this is always a threat and always something that can happen.
So how can we protect against this? Um there is some work to be done in the model. You want the model to understand as much as possible what is an injection, what is a true instruction and what comes from user unsafe context. Um, obviously you want also want to restrict the privileges of your agents. You want to make sure that at runtime they can only do what they need to do and this can be dynamic depending on the step depending on which workflow you could change the permission uh dynamically as well.
Obviously send boxes are very important here. Um and then you have runtime guardrails. You have everything that you uh need to observe the behavior of the agent itself by looking at the traces the low-level traces. And that's where sandboxes are also so important. You can look at each system call at each network call and from this you can derive heristics and understand what sandbox is actually trying to do. If the agent is following a nominal pattern or doing something that is a bit abnormal.
So the problem with this is when you add all those security guard rails everywhere, you basically break all the use cases because you had such a powerful agent, such powerful tools, but then now you're restricting everything because you fear everything. And so this is a clear dilemma that we have to deal with. So we are solving this with dynamic collaboration and making sure that as I said you will grant the right amount of permissions at the right moment when you need it and when you have a good reason to grant them.
Um and also making sure that we can filter which operations should be allowed and why. This is a very quick demo of a typical prompt injection where we will ask VIP CLI to perform a task by looking at a GitHub issue. So it's going to go very fast, but by looking at the GitHub issue that has a prompt injection in it, it will actually post a command here that exfiltrates a secret key just because you took the uh issue content that was posted by an adversarial user.
So we're working on prompting injection rules where you can describe a set of classifiers. You can pick which one you want. It could be PIIs, it could be secrets, it could be uh you know company knowledge and things like this. you define a source, you define thresholds, you define the kind of moderation that you want to use. Um, and you can define as well what you want to do uh after that. So once you've defined your rule, you will now attach it to uh a policy. >> And the policy will allow you to basically find out where you want to apply this rule on which product surface because it could be vibe on the web, it could be the mobile app, it could be the CLI.
Uh so this is exactly what the rule allows you to do as an enterprise admin and you can define the threshold here. You can apply it and then you just need to enable it and now you will uh use the VIP CLI uh again perform the same action but it will be stopped at the network level without the user having to configure anything. And that's the power of enterprise controls is that the user never opted in to get protected.
This is a policy that has been defined and pushed at the enterprise level. And this is the kind of work that we're doing on uh on safety. But safety does not is not enough if your AI is not sovereign. So what is what is sovereign? Well, you have data sovereignty. Where is your data? Where uh or your users data, your customer's data, where is it stored? Where is it retained? Who has access to it? Which hyperscaler or which company because of jurisdiction or something could actually read this data? intelligent sovereignty as well is when you analyze the behavior of the model when you try to tweak it when you finetun it where is this happening um compute as well so now you have all the inference when you do inference you all by ZDR everywhere but still the data transiting through some places uh and understanding this is also so important or is operational because you need to know exactly what's going on in your company uh exactly what's going on in terms of security uh audit logs and and things like this.
So the way we answer this is typically this kind of data center that we are currently building. Mrol was just a model company some time ago then it became a product company and now we are building our own data center. Um this is typically leis uh 10 megawatt um facility uh near Paris with B300 from Nvidia. So this is one of the largest uh data center with this kind of cutting edge hardware uh from Nvidia in Europe. Uh and this is typically where we serve our critical workloads internal ones.
This is where Lat might be training right now. But this is also where we put the most sensitive customers uh workloads. And this is how we can have from end to end a full sovereignty path uh that represent to the customers including agentic safety but also including inference and training. But you also need models obviously for uh all of this to work. Uh this is basically the family of models that we have the generalist ones.
Uh different size for different matters. I know you're all expecting a big one uh and I hope we can deliver it very soon. Uh but we've also been focusing a lot on specialist models. There are two kind of specialist models. The specialist models that serve many purposes. We're going to talk about OCR and VR in a moment. But also all the models that we fine-tune for specific customers that have specific set of data. So think about the rare languages, think about you know some companies have their own programming languages that nobody has heard of.
So obviously by default coding models are not good at it. So how you can actually um you know sit down with a company and make a specialized model makes an amazing different. Uh and we have a few experimenting uh models that allow us to uh refine the frontier of science. Uh linstrol for formal proof. Uh chillstrol we announced a month ago. Uh you know all talking about jev uh everywhere but chill troll is basically you know answering a moderation question extremely fast uh with a very nice architecture uh for safety purposes.
If I dive in a bit into the um into the specialist models, let me talk a bit about Voxrol uh our family of audio models. So we do have uh texttospech models and for this I will just showcase uh I would say something that I said last year at AI engineer um and you tell me which is what we really saw that the mission of providing open source models to actually uh bring the AI to the masses uh was also compatible with pushing enterprise beyond and building some strong um focus.
So what we build at mistrol uh so once you actually collect all this data you want to watch people do what they do um and in enterprise context it's really important to use this to understand how people are actually doing their job how are they processing their days what kind of workflows are they putting in place so as you might have guessed uh one of them was uh AI generated the other one was me uh so this is the kind of voice cloning capability that we are heavily investing on to make uh this kind of models extremely useful.
The second sample was not me at all. Um and and this is the kind of capabilities that we're working on for uh speech to text. We work on higher accuracy. Um we make sure that we can uh uh yeah basically mimic any kind of voice, any kind of accent so that it can be useful in any kind of settings. For the speech to text, uh we already have state-of-the-art models and we're improving them a lot. Um what we're working on is real-time derization so that you can have in a meeting a live transcription but also knowing which speaker is which.
Uh making sure that we support more languages obviously accurate time stamps uh more languages uh making sure that it works in very noisy settings. Uh and we can do live transcription actually. >> So within banking uh we have a very important process which is called QYC. It's a long process. So with Mistral AI, we started to work on a specific QYC process for corporates in Belgium and we developed two agents. We were able to go from 80% of files incomplete sent to the middle office to only 10%.
For the clients, the application takes much. So it was no cheating. It was an actual live transcription and live translation uh of a video using our uh latest models. So we are actively working on making it more useful for customers. Uh and this is something that is key in the way we work. We're not trying to you know benchmark everything. We're trying to make things that are going to make a difference on the ground with our own customers.
Briefly about OCR 4.1 the model that we announced uh this summer. So this is also a state-of-the-art uh OIR model that is now able to transcribe any kind of handwriting, any kind of scanned old document from the 60s or something. Um I went uh slightly fast here. So this is typically a pretty hard document where you can see the labling. Uh you you distinguish captions, titles, text, tables, um images as well. You're able to extract all of this.
Uh this is a very powerful model that is used uh both by our typically our document library. So document library within vibe is where you will upload a trove of data. It could be you know terabytes of data uh unstructured documents or anything and everything will go through the otr pipeline if needed so that you can actually query it um at runtime whenever you need it. This is exactly what we use for advanced rag uh pipelines with our own customers uh when you have this kind of uh of of setting.
So we're working hard to make it more robust to understand more languages as well making sure that it is very very noise resilient. Uh but this is the kind of uh you know models that are state-of-the-art um and that we are actively working on because they actually solve enterprise issues and they're I mean they help building true use cases uh that we can implement that are in production uh and that change a lot of things for uh our own customers.
Uh those are the playgrounds. So you can go to the studio, you can actually test those models. Uh you can test the oops sorry uh you can test the uh voice models, you can test the OCR models uh with the bounding boxes that I didn't mention but that we added uh a months ago as well. Uh and that allow you to do extremely advanced use cases based on all of this. So briefly we're working on a new vision of work trying to go beyond the chat assistance to make something useful and merge the personas between the end user and the builder.
Uh we're working hard on agentic safety making sure that we can keep very uh powerful agents uh but still understand what they're doing making sure that we can protect uh company knowledge enterprise knowledge and that we have the right set of primitives um we are working on sovereignty a lot making sure that we give the right level of control to all our customers investing a lot in data centers I mentioned one in France but also all across Europe so that we can be a fully full stack uh AI company and we're still working hard on advanced model.
Uh the science team has never been bigger. Uh there are many things happening on those different fronts. Uh obviously the LLMs remain one of the major use cases that we're actively working on at the moment. Uh and I hope we have uh very interesting things to share uh very soon. Uh thank you so much for being here. Uh I hope you enjoyed the summit. Many people in the team will be here. Uh some of them have uh aren't shirt but otherwise feel free to reach out.
Remember that uh we are 1300 but still hiring. So uh also come and say hi and thanks for all the speakers. I want to thank also Jen Aliser and all the Debell team for organizing this and obviously Ralph uh for being our MC tonight. Thank you so much and enjoy. Nice. All right. Thank you, Lilio. Let's give it up one more time for Lilio. Yeah. [applause] Yeah. Absolutely. I think uh it was on on point. I think we should give it up also for Jen uh Ali's there and all the ML team for putting this together.
Okay, so I have one sad news. The sad news is that yeah, tonight is over, right? You can go back and relax, etc. But I also have two better news. Okay, so we're going to we're going to come back tomorrow. There are just so many things that we're going to cover tomorrow. You're going to hear from Black Forest Labs. You're going to hear from Nvidia, from so many folks. So, the lineup is going to be extraordinary tomorrow.
All right. So, uh who's coming back tomorrow, by the way? I see all of you. Yeah. Okay. So, we'll see you tomorrow. But for tonight, we have a reception at the expo. Okay. So, we can all just slowly go there and um just enjoy, meet people, take the time and and talk to everybody. Some of our speakers are going to be there, some of our sponsors are going to be there. So, please join us all upstairs at the expo and see you in a minute.
All right. Thank you so much, guys. You are awesome today. Super.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.