Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
Most replayed moment at 6:57
5.9x that video's typical replay level
do light mode. It's I It's not my nature, but sometimes. That's better, yeah? Okay. So we have we have a model and we're trying an old LG Sorry. We We shouldn't have seen that. No, we'll
Said at 6:50
The graph counts replays. It does not show where viewers stopped watching.
Words
4,177
Runtime
22:31
Speaking pace
186wpm
Reading time
17min
186 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
[music] >> Thank you all for coming today. I really appreciate it. Thank you for coming to this talk which is Black Rice Labs Flux Open Research and the future of visual AI. I'm going to start quickly with a quick intro of myself. I'm Stefan Batyfull, sorry. I'm a developer relations engineer at BFL. And I want to start with two questions. First, who here knows about BFL? Raise your hand. Okay, who here knows Flux? Okay, about the same people actually. But for the people that don't know BFL, I have
93 words, the words spoken in the first 30 seconds at 186 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 296 |
| Average words per sentence | 14.1 |
| Longest sentence | 67 words |
| Questions asked | 17 |
| Sentences containing a number | 25 |
Most used terms
Filler phrases
192 in total: you know 63 · like 62 · actually 32 · uh 17 · um 8 · basically 5 · I mean 2 · kind of 2 · literally 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
[music] >> Thank you all for coming today. I really appreciate it. Thank you for coming to this talk which is Black Rice Labs Flux Open Research and the future of visual AI. I'm going to start quickly with a quick intro of myself. I'm Stefan Batyfull, sorry. I'm a developer relations engineer at BFL. And I want to start with two questions. First, who here knows about BFL? Raise your hand. Okay, who here knows Flux? Okay, about the same people actually.
But for the people that don't know BFL, I have a quick intro so you are not lost. But BFL at a glance, we are the team behind Stable Diffusion, Latent Diffusion, and the Flux models as well. Our team has more than 200,000 academic citations. And we don't only build models, we actually also work with enterprises and customers with them. So some of our customers are Microsoft, Adobe, Canva, Mistral, and many more. And the way we started is we started in August 2024 with Flux 1.
Flux 1 was the first breakthrough, you know, that was the model that was the big competitor to Stable Diffusion back then and that was really the breakthrough where people were like, "Oh, this is a really cool model." We released it in open source in the first place. So really that was the one that was text-to-image only and you could run it on your laptop. That was a game-changer as well and the anatomy was really good in comparison to the other models and especially in comparison to other models that were way bigger.
So this is where we really had a breakthrough and Clem from Hugging Face actually gave us a shout-out back then. This is fairly old but Flux was actually the model that was the most liked on Hugging Face back then. This is not true anymore, uh but back then that was really the thing and that was really good and really big actually for a company that was, you know, coming out of nowhere and just released this model. We then released Flux Context, which was the first open-source editing model in the world that was like the combination of text to image and image editing as well.
This one Now, what I'm showing you here, what I'm showing you here, you know, it's obvious because now we have editing models everywhere, but back then that was a big breakthrough where you could do both text to image and image generation at the same time. If I have an example here, we have this input image and then you remove this snowflake from the face, you know, you can see have the character consistency, but you can also then move this person to be in Freiburg, which is where our headquarters is.
You know, she's chilling, uh taking a selfie in the streets of Freiburg, and then you can, you know, do some local editing where you can change the background to be snowy, and then you also have snow on her face and everything. It was also the model that was really one of the fastest back then. You know, if you remember, this is the time where you had the first GPT image where it would take like 40, 50 seconds to generate or edit images.
Whereas Context, if I remember correctly, was like 7 to 8 seconds. It was also really useful to tell stories. I've seen lot of use cases from our partners, from our customers where they would start with an input image and then create a storyboard like we see here. We have the famous seagull here, you know, which has the VR headsets drinking a beer in a bar. But then you can actually create other things. You can have a friend that is joining and that is then drinking with them.
Now, you know, I guess they got a bit tipsy and they're wearing hats in the bar. Then they're going outside. And then, you know, this is the story you could create and that was really useful actually for video model or for animation models. You know, you would give those images as input frames or as end frames and then the video model could then create different content. In November, we released Flux 2, which is our steps towards what we call visual intelligence.
Flux 2 was uh as our base still our base our best model, sorry. Um those are samples, which I don't know if you can see clearly, but in my opinion, they are like really amazing samples and it's impossible to basically tell they are AI-generated. If you look at the hands, if you look at the veins, if you look, you know, at the bracelets of the person on the left, there's personally no way that I would tell it's AI-generated.
Same for the turtles you see on the right side, same for the dog or cat actually there's in the bath, that would be very very hard to get the sample like this. Um but those are AI-generated. And then you have more as well, so it's not only the people or animals, you know, you could do like some proper product photography. You can see it with a waffle here on the bottom right, or you can make some very cute images like this person, you know, on the left that is really driving the moped with some balloons.
And this is what we released in November, but it's not only an image generation, it's also an image editing model at the same time. Um you can see on the left we have six images that we give to the model and my prompt was literally like create an outfit with those images and then the model is intelligent enough to actually, you know, make things that make sense. Like the jacket, you know, is worn properly, same for the tie.
And on the right side, it's a bit more of a simple use case where you have the sofa and then you have to imagine, you know, maybe you are an e-commerce website or you're sofa maker and then you want to people to imagine, you know, what it looks like in your flat or what really it would look like if you had to buy it. And those use cases are really really important and those are like the main use cases we have currently for Flux 2.
But it also takes, yeah, up to 10 images simultaneously, so you can really edit a lot of images at the same time, and then you can create some magic things. It's very good at character, product, and style consistency. And what I want to make clear is that BFL as a company and as a research lab, our first operating principle is to release state-of-the-art models. This is what we want to focus on. This is what we want to do as a company, you know, we want to raise the bar on quality with every release we do.
So, we did it in the past, you know, with Flux 1 when it came out. We did it with context. Our Flux 2 was our best image model to date. It's the first one we released that was actually multi-reference as well. It was state-of-the-art in the open-source world. So, it was state-of-the-art for text to image and image editing. In January, we released Flux 2 Klein, which is a step towards like interactive editing and interactive image generation.
It generates and edits images in less than a second. I'll talk a bit more about it later on during the talk. But I think the fastest it can do, if I remember correctly, it's 500 milliseconds for editing and 300 milliseconds for generation. So, basically real time. But this is not it. We also have more things that are coming, and this is where I want to talk about today. So, I mentioned it. We are a research company first.
We publish things in the open. We publish paper. We really want to make sure that the field is moving forward with us, you know. This is our big focus as well. So, it's state-of-the-art model, publishing things in the open, and that's what we want to do. But I want to first take a step back and tell you a bit, you know, about like how do you train models, and especially generate um models that are generating content, generating images, and everything.
You you when they generate things, when you train them, they actually don't understand what they're generating, you know, they don't understand that my glass here should be actually on the stable. I shouldn't go through it. Because you train them, you have images, and then you know, you add some random noise to those images, and then you just try to denoise them. That's what you do. That's what those models are doing.
And when you denoise images, you never learn, you know, that my glass shouldn't go through here. You didn't never learn that, you know, you sit on the chair, you shouldn't go through it. So, what do you do? Is that you use you do what is called like representation alignment. So, you use an external model that actually knows about this, and that is an encoder that is like an image encoder that is teaching our model, "Hey, he is currently sitting on the chair.
He shouldn't go through it." And those models are external, and they are really like trained to segment images, whereas our models are trained to generate images, or generate videos, or generate audio. And you try to align them to be on the same objective, so that our generative model actually learns, "Okay, you shouldn't go through the chair. My glass should stay on this table." And this is great, because it really improves the way generative models are working.
We can see here on the right, you know, it is 70 times faster to actually converge and to reduce the loss when you use this external alignment. So, you're like, "Okay, this is great." But as usual, if something is working well, there are also counterparts to it. So, the first one is that you have a scaling ceiling. You imagine you're working with a model that is external, that has been trained at a checkpoint, you're not changing it anymore.
What if you train a new model, and you have a generative model that you want to scale up? You're still like limited by this encoder that you have on the side, you know? You never actually scaling up fully with that. Also, those are specialized in modalities. You have an encoder, for example, DinoV2 and the other one that I can't remember. It's specialized in images only. What if you want your model to generate images, audio, video, and more?
You would have to have encoders for all of those, and you can imagine then you would have like a very Frankenstein setup, you know, nothing would really make sense. And the objectives also misalign, so I've said it before. We want to generate content. We want to generate images or audio. The other one is here to segment things. So, like they have different objectives and you're trying to make them work together. And it works great, but it's also not perfect.
Here, you see on the right side, we have DinoV2 and DinoV3. DinoV3 is a better model technically, per se, than DinoV2. But when you train your model, actually getting worse performances, you know, DinoV3 is here in right in red and green. And so, you're getting worse performances. So, you're like, "Okay, like this is supposed to be a better model. And yet, when I do train a model to generate things, then it gets worse." And there's also like not really any rules as to why, you know, certain encoders should work or otherwise shouldn't.
So, how can we solve this? How can you teach, you know, a model representation directly without this external encoder? This is what we released about a month and a half ago now, which is a research paper. You can read it. It's called Self Flow. It's in the open. We released it to really make sure, you know, we're moving the field forward again and it's not only us benefiting from it. And it's basically a scalable approach to training multimodal generative models.
So, they use self-supervised learning, so you don't need any other models, you know, to train it. And I'm going to try to go in tiny bit more details into it. But we combine representation learning and generation in the same flow. And you see here on the left, you have videos, images, or audio. You have different modalities. What do you do when you usually train a model? You add some noise, you add some random noise, you try to denoise it, and then you align it with the encoder, you know.
How do we do it then? We actually add two different kind of noises that are both random and they're both different. The first one we're adding is actually we're adding a lot of noise to the assets. So, this is the one you see at the top. And the other one we're adding like a low amount of noise. This is what you see at the bottom. And the idea is that then we have two models that are actually working together. We have the student one, which is always getting the images with the most noises and is trying to denoise them.
And then the teacher one, which is basically a more stable version of the students, is always getting the low um noises images. And then the student one is actually trying to learn two things at the same time. It's trying to minimize the loss for the generation and the loss in representation. And this is how then you actually work across different modalities. You know, this is then you only have one model, you don't have anything external, and if you actually scale up your model, then you're scaling up your student, you're scaling up your teacher.
And you don't have to worry about the encoder that you have on the side anymore. And this is where we're working on. This is something, you know, we are currently using for different models that we're training. Um and this is, you know, where we believe the future is going to be and to get rid of those encoders that we have. We actually trained models. Uh so, those Disclaimer, those are research models, they're not meant to be released in production.
Uh but we released actually one model on all those modalities. On the left, we're comparing flow matching, which is the usual way of training models, with ours. And you can see we are better in audio. So, this is what you see in the orange on the right side. And then, we also better in images. So, the dash lines is the baseline. And then, we are like the full line where we can see we're also better at images. And also better at video.
So, with this approach, without having the encoder and the external model that you may struggle with, you actually get better at every modality that you're training your model on. It's also converging faster. You can see on the right, you know, the baseline is converging is actually hitting a plateau. Whereas, we are converging faster. And we're still, you know, decreasing the loss. And I'm pretty sure that if we were to go towards 2 million steps, you know, the baseline would really plateau and then wouldn't really get any better.
Maybe actually get worse. Whereas, we would still go down in loss. And this is the difference between the two. If you used Flux in the past or if you used different models, you know, to generate, you may have noticed the text might not be perfect or, you know, things don't really make sense. This is what you see at the top. Where on the left it's like the future is Flux. But you can see, you know, you have like some letters that are missing or maybe you have two letters like on words, for example, you have two L's instead of one.
Whereas, with this approach now, in the Cell Flow approach, you can see at the bottom everything makes sense. There is like they learn representation, you know, they learn that Flux then for the letters should be like one next to to the other. And the same on the mirror, same on the tree. And this is where we believe this is the future. But on top of this, we can see some comparisons here. On the left is the baseline.
Where again, the letters are wrong. Um on the right, you can see that the letters are correct. Here is the same for the anatomy. Where you see on the left, you know, you have like a face that's looking a bit odd. Let's put it that way. And on the right, this is the one with Cell Flow. And again, this is not like a production model where, you know, you expect the face to be like perfect, but you can see that the anatomy is way better than what you have on the left.
What I want to show you as well some different generation if it loads. Yes, thank you. Uh this is also possible This is also possible for video generation. So, this is the same model that has been trained on images. Now also can generate videos. On the left, you see the baseline. So, weird way to do a push-up. Let's put it that way. Whereas on the right, it's a perfect form. You know, the arms are correct. The hair as well is correct and nothing is wrong with it.
And this is, you know, a way to actually fix all those artifacts that you may see usually in generations. The same here for the birds. Oops. My bad. Yes, thank you. Uh it's the same here for the birds where Thank you. Where you see on the left side, you have the baseline. There's a lot of flickering. There's a lot of like, you know, weird things happening because the model was using this encoder was trying to align things.
Whereas on the right side, we sell flurry just does it perfectly and like the the bird, you know, is walking on the floor and there's no flickering or anything. But, it's not only about images or videos or audio. You train those jointly. So, you can also actually generate things jointly. We have here an example of a video and audio sample where the idea is that we have someone that is saying "Hello from the Black Forest." Uh I will just play them and you will hear the difference.
Again, this is not a production-ready model. So, it's not like perfect, but you can hear the difference hopefully. Hello from the Black Forest forever. Hello from the Black Forest. Hello from the Black Forest forever. So, this one was the baseline where if you hear it correct if you try to pay attention to what he's saying, you hear like "Hello from the Black Forest." There's a bit of like, weird things at the end. Hello from the Black Forest for a first.
Hello from the Black Forest. Whereas on the right side you can see, you know, the prompt is really just say hello from the Black Forest and then it ends here. And yeah, this is the same model that was trained on those images that we've seen before on video and on video and audio. Hello from the Black Forest for But this is cool and this is great. But what if you could also teach robots on how to use this? This is also the same model.
This one is trained on actions and not only on images, video, or audio. So it can also predict actions. And what I'm going to show you now, it's a robot that is trying to pick up a can and make it closer to us. On the left, this is a baseline. Again, you see some like flickering. You see like the the arm is doing weird things. Whereas on the right for the same amount of steps, you can see how the robot is picking up the arm directly and like bringing it closer.
And this is where we're going as well as a company. This is where we're really interested is like not only image generation or video, but it's also doing actions and doing more things toward physical AI. And there's more. It's also how do we make our models faster? Because this is really important for us. This is a demo of Klein, which is, you know, like near real-time editing. You see it on the right side. This is, you know, generated with Klein on Korea, where you see the edits and this is not a video model.
This is those are images that are always editing in real-time. Not only they are faster, they also actually, at least on par or better than other models. And I'm almost out of time. Oh, they added 5 minutes. So I don't know if I'm Okay, cool. So I'm not out of time. So I can chill. Uh, so yes, here on the left we can see, you know, we have Klein uh that is 4B and 9B that is compared to the other open source models. So it's at least on par, while the latency, you know, it's like 0.5 seconds, while Kwen is like around like 15 seconds, you know?
And if you are like on par and you're like way faster, then this is really, really good for us. Same for imagery image, you can see the editing. For Klein 9B, we are at like a tiny bit more than 0.5 seconds, whereas Kwen is still at around 15 seconds. And then same for multi ref, you add it, but we're still at less than a second, whereas Kwen is more towards the 20 seconds. And this is what is really, really important for us because you really want to actually generate things in real time.
This is where we believe the world will to visual intelligence. This is where we're going as a company in the future. And why does it matter? As like I mentioned it, real-time generation, so you can imagine you render mock-ups as fast as you think. You know, you don't have to wait, you don't have to wait like 10 seconds, 20, a minute or two. You do things in real time and you can guide in real time. This is also where we're going.
You can think, you know, interactive visual engines for gaming or films, where you really render a movie as you're prompted. On top of this, there's also world models. The idea of world models and behind it and why it matters for us, it's you train your models to understand and simulate geometry, relationship and like different interaction of the world. And you may be like, "Okay, that's cool from a research perspective.
Why do we care?" The reason is robots. Uh that's why we care. That's why, you know, robotics and automation, this is where we're taking BFL, and that's why we also want to go towards world models, is to train agents in those generative world to scale self-driving and automate every manufacturing. And I think that is it. Thank you very much. >> [applause] >> Do we take Yeah, I think we can take questions. Yeah, I have something like Uh you mean the data?
Can't really. Uh this trade secret. I mean data is very sensitive as you can imagine. So, I can't really share this. We're partnering with a lot of people though for it. Mhm. Well, this is what the model is learning. It's basically like those representation, you know, the model is learning that in itself as like the state and it has like some kind of memory. And this is uh the way we do it. No, it's Yeah, it's the context window that you have, you know, you train and then you have the tokens and then they'll be like, "Oh, look, I've moved.
Here is where I should be then next." Uh define long. What do you call with long? Indefinitely. Uh that I'm not sure. I mean, there's always going to be a limit. Uh so, you may have, you know, like a sliding window. Um but this is the way we see it. Thank you. >> [music]
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.