Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
Most replayed moment at 13:49
2.3x that video's typical replay level
Fable uh and it runs into an unknown, ask it to log it, right? So that um you uh you can see where the deviations happened and then you can sort of figure out why as well, you know? It will usually give you some context about what happened.
Said at 13:43
The graph counts replays. It does not show where viewers stopped watching.
Words
11,942
Runtime
1:17:14
Speaking pace
155wpm
Reading time
50min
155 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
[music] >> So, good morning everyone. Thank you for being here so early in the morning and to be able to be the few ones who pass security to to be actually here. Um I'm I'm Guillaume and I will talk tell you about uh Gemini in general. Um Like so this is this is my life as as Nano Banana sees it. Basically, I I joined Google 6 years ago. I used to be a video game
78 words, the words spoken in the first 30 seconds at 155 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 638 |
| Average words per sentence | 18.7 |
| Longest sentence | 94 words |
| Questions asked | 28 |
| Sentences containing a number | 46 |
Most used terms
Filler phrases
675 in total: uh 268 · like 145 · um 144 · basically 56 · actually 29 · kind of 26 · you know 6 · I mean 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
[music] >> So, good morning everyone. Thank you for being here so early in the morning and to be able to be the few ones who pass security to to be actually here. Um I'm I'm Guillaume and I will talk tell you about uh Gemini in general. Um Like so this is this is my life as as Nano Banana sees it. Basically, I I joined Google 6 years ago. I used to be a video game producer before. I initially worked on Stadia, the streaming game companies that product [clears throat] that's uh that we killed as so many other products.
And and I've been at DeepMind for 2 years now and as I worked as a what we called a developer advocate. So if you're not familiar with what the developer advocate, basically my my job is to make sure that when whenever we release things or whatever we release, you guys the developers everything you need to work with our products. So you you need documentations, you need code samples, you need >> [clears throat] >> demos.
Now there are new things that you need as well like skills and prompt guides and and things like that. So making sure that you can you can start right away and it works. Um and on the other on the other side when I'm talking internally, I'm that's the reason why it's advocate because I'm advocating for the developers and trying to bring some let's say common sense to the internal teams and making sure that what we what we release makes sense in the in the real world and it's not entirely developed for Google by by Google.
Um, a very good example of that is the Imagine models. Uh, when we release when we had like Imagine and Nano Banana, the each model has its as as its own set of API. It doesn't make any damn sense. Like you like a normal developer should be able to just swap the model name and it works. Uh, and I've I've I've fought quite a long time for for that. I never managed to win, but in the end I think the Imagine brand is doesn't exist anymore.
So, I kind of win by default. Um, but yeah, you see you see the kind of the kind of work I I have to do on a daily basis. Um, as I said, I'm mostly working on the Gen Media model. So, we like and that's why the talk is about Gen Media. So, uh, if you if you look at the world, your phone, and anything like medias are everywhere. So, you have like images, videos, sounds everywhere. So, it's really cool in our world. And that was at the core of what DeepMind is is building with with with our models.
Um, we like you you you're starting to to hear with with Lukan's new startup about world model, but that's exactly what we have been trying to build at Google for since the beginning. Um, and our vision of world model is that it's it's a model that as its name imply understands the world, but meaning it can ingest as many modalities as possible. So, sound, videos, audio, sensors, whatever kind of all five senses. And then um, uh, outputs or like uh, yeah, talks in in different as many different modalities as possible.
So, audio, text, and and and so on and much more in the in the future. Um, And we we tend to to have like specific models. We have all image generation models. We have all video generation models. But deep down the the goal is really to have like one model that encompass all of that. It's just that for release purposes, it's easier to ship specific models than to only always update the main model and like have risks of breaking something else at the same time.
But like a quick a quick story. Like when we released Gemini 1.0, it was 2 years ago. It's It seems It feels like it was like years ago like eons ago. But the first Gemini model Gemini 1.1 was meant to be multimodal because all of our models have always been multimodal. But since I guess the testing was not finished or whatever whatever reason I wasn't there yet. They they removed the image understanding the music model understanding inputs from the model.
So the 1.1 was not multimodal and then 1.5 came and then this one was multimodal in and I think that was the first one that was that was doing that. And which once again is crazy like it's was year and alpha go and then now you don't imagine working with model that is not a multimodal one. But still when you were using that that one very often you were giving it an image and it was telling you oh I'm sorry. I'm just a LLM.
I can't I can't deal with images. And that just because some of the training that was made that was added at at the end of 1.0 that was oh you know how to deal with image but if you are asked don't use it. Though some of that training was still remaining in 1.5 and coming up for a from time to time. So that was yeah. That was kind of annoying for until we we ship we switch up to to 2.0. So as I said we don't only have the Gemini models in a DeepMind, so we have all of the image, video, music generations.
We have a couple of other very specific models. We have the robotics one that is kind of a multimodal one as well because the vision is very important for for robotics. Uh we have a bunch of agents that we are shipping. I think this year is going to be the year of agent, the real one. Last year was everybody talking of agents. This year is the is a year where we are actually going to build agents. And we have the open models and I forgot to update the slide because it's not Gemma 3 anymore, it's Gemma 4 since last week.
And then bunch of uh research key models like AlphaGo, Alpha Alpha Genome, the Waze or next and so on that's very like very specific for research purposes. Um just on the on on Gen Media, we've shipped things uh on average more than every month. Uh on on the whole, if you just take all of DeepMind, we're shipping things every 5 days on average. Some Some weeks we're shipping two or three things. And if we add all of the small small features, we're shipping multiple times per week, so it's uh that's why I'm I and most of my colleagues are very busy because we always have like new things to to document and talk about.
Um So very quickly about the new things that we shipped recently. Uh Nano Banana that we shipped Nano Banana 2 uh which has uh new uh aspect ratios uh from uh off uh like 520 pixels to to 4K. Um it has search grounding as as you all know, but it also have like image grounding. So if you can ask it to search for images on the web and use those images as grounding so that it knows uh uh like yeah, it has better knowledge about what what things look like.
It's very useful for uh architecture stuff, for example, for uh animals and and things like that. Uh, there's a lot of rules in total that I I was not able to to all figure out, but for example, buildings, they have to be old enough, otherwise there are legal reasons we can't use the images. So, um But but it's still uh very useful when when it works. Uh, we have the video models. You've heard about video of three, video of 3.1.
We um we just released last week uh video of 3.1 light, which is our cheapest model. Uh, I think it's 5 cents per second, so 40 40 cents for for one video, which is very cheap. Uh, and the idea is that you can iterate on um on on the prompt that way and then upscale afterwards. Um And we have Lyria, which is our music generation model that we um that we released uh 2 weeks ago. Um and with which you can create either 30-second clips or full songs of 3 minutes.
I will I will show you demos of all of that afterwards. That's the point of the of the session anyway. And uh but I there's also another Lyria model that uh people don't know about that is called Lyria real time and that's actually my my favorite model. Um And this one is basically you you you create music as well, but it's not a uh a diffusion model like the others where you just give a prompt and you get something out of it.
It's uh it's a predict model, so which means it's uh it's it's basically a live model. So, it creates music and it continue to create music uh in real time until you stop it. And you can just send you prompts and it will like a DJ mix and like swap to the to the new music in real time and ever real time 2 seconds. Uh, but uh but then uh and that's that's quite fun to play with. Uh if we have time at the end I will I will show you a demo.
Um but as I said, uh this is meant to be a workshop, so the goal is for you to to play with with the models and for me to show you codes instead of slides, so um this is the content that we're going to use. Uh that's this one as well, if you can read what I wrote. Um And if that doesn't work, just tell me. Are you all in? No. >> Not. >> No, it doesn't work? Oops. Let me Let me try to check that the link works. It's working.
Yeah. >> I know. I had a typo. >> Cool. Um are you all in? Okay. So, uh full disclosure before we start, uh this is using Gen media models, which means they are all paid models. So, uh running the the notebook is going to cost you something like $1. Uh and you can just skip the video generation because that's a most of the of the cost of that of that $1. So, um the Yep. Uh Can zoom in a bit. >> [snorts] >> Um So, the the goal of that uh that content I I is to illustrate a book using the different Gen Media models.
So, the what we're going to do is that we're going to to take a book that is is an open source one that I took on the Gutenberg library online library. And basically, we are going to use Gemini to come up with prompts and then the Gen Media to create the content for the prompts so that we can we will have images of the characters, imaging of the scenes, videos of those those things and so on. Um and and that's that example comes from what we call the cookbook.
So, we have we have this GitHub repo [clears throat] where we are posting examples of like both quick start guides to explain how to use a new model or to use a new feature and also example like full fledged examples like this one on how to go further and to mix different features into into content. So, if you if you're looking for ideas, that's that's a good place to to check what you can do with the models. So, basically, introduction building.
So, let's start. So, the first thing is that you need to install the SDK. Um you need the latest one because of music generation that was shipped last week uh 2 weeks before. So, actually, you need the next one. Um So, it's starting. And I should have done that. Um and you will need an API key if you don't have an API key from from AI Studio. And you need the other said you need that API key to be a paid one. Um if you don't have a paid API key, you can still use the image generation examples by using the Nano Banana one model, which has a free tier.
So, you can just swap the model we're going to use for for this one. And why is that so slow? I think. I think it's good. Okay. So, then I'm just loading the the API key and I'm creating the clients the GenAI clients and you can see there I I did these parts that you we don't add in all of our examples that is basically the auto retry thing because like if we're using Nano Banana 2 at the moment so the especially in the evening when the US wakes up, the model can be overloaded.
So, that's that's basically that part basically says that it's going to be automatically retrying five times after 2 secs. Um a bunch of imports and then we're going to select all of the models that we're going to use. So, in this case 3.1 flash image preview which is Nano Banana 2, Gemini 2.1 2.5 flash no. Let's use 3.3 flash instead. Lyria clip to create 30-second music clips and um and the TTS model this one the pro one.
Up. So, and I added this checkbox that should have been false by default but I made a mistake yesterday evening just so that you can if you you won't be able you should not be able to run the notebook by mistake if you don't want to pay. Um then just I'm just setting limits here because like a book can be can have a a of chapters lots of characters and that can be quite long to to to generate. So, that's why I set up limits on how many how many we want for the demo purposes and cost purposes as well.
So, as I said, we are going to use an open-source book that's will the wind in the willows from Kenneth Grahame that I don't remember reading, but I think in the UK it's quite well known. Um So, basically, I'm just downloading it from the Gutenberg project and uh and here I'm using uh client file upload, which is you might know that we have basically two ways of using the Gemini the the Gemini models. We have the AI Studio Gemini API way and you have Vertex.
And the main difference between the two, and I can maybe that's a good time to switch to this slide. Up. Um But uh we you you know at Google we like to basically create multiple products that are doing the same thing and confuse our users. That's kind of our uh motto. Um so, that's we're doing the same with with AI. So, we have a bunch of but actually when you think about it, it's it can make more sense. So, on the on the left, we have what we call the consumers app.
So, it's uh fully developed uh app for for that are easy to use for anybody who is not technical. So, you can do plenty of things with the with Gemini, but you you are you are lacking as a developer you might be frustrated because you're lacking control about which models, which features, which parameters are being used. And on the other end, we have uh Vertex AI, which is uh the exact opposite. It's meant for enterprise, so you have a lot of control.
You can control on which data center it runs. You can uh you have control over about your buckets, with access to what and and so on. The only thing is that it comes with like that with great powers comes great responsibilities. So you it's it can be a pain for people to start there. So that's why we have the Gemini developer API that are kind of a middle ground for for developers where you can just create an API key and then start using the um the models right away.
We which comes with us security risks as well that if your key leaks every anybody can can just use it. Um Um And and AI Studio is like is is a kind of the the same vein that it's it's meant to be a place where you can test the model and and play with them as easily as possible. And the good thing is that we have the same SDK between Vertex AI and developer API so you can swap from another one to another. So there's no wrong place to start playing with the models because you can always change.
Um and then we can skip that. And I can go back here because actually what I was going to say is that we actually have a few differences between when you're using Vertex and when you're using the Gemini API. Um because like the the goal of the Gemini API is basically to hide all of the complexity from Vertex. And one of those complexity is creating buckets, creating ACLs for the buckets, giving rights and and and all of those things.
So the Gemini API of us that that API that calls that is called file upload. And basically what it does is that you upload a file and then it's easily accessible from the model. Um So I upload the file and then I I will use I will use what we call chat mode. So basically what it does is that it's it sends request and it keeps the history of it so that it's it's easier to to keep all of the context. And in this case it's going to be good because we are going to feed the whole book to to the model thanks to the the large context window and that's why it's going to be useful in in in our case because I'm going to run that while I talk.
Up. Um it's it's going to be useful because we like for image generation it's it's always a good idea to know what's what have been generated pre previously so it it keeps the same styles and and things like that. Um I'm also going to use structured to the output so that we like we have a structured output what the year the model outputs and we we can we can talk the same language which is going to be very simple. I want it to generate prompts and I want the prompts to have a name so that we know it's if it's a chapter or a character or something and the prompting itself.
So I'm just creating like initializing the the chat clients with the response type JSON the scheme that I that I'm providing and something that we shipped yesterday that is service tier priority. So don't do that yourself. I think you should remove that line. Um What what it does is that it basically it's it's yeah we shipped that last week. So we have uh three service tiers. You have the normal one. You're paying the normal price.
You're like in the queue with everybody else. And we have another one that is called flex and that's basically I don't care if that takes a long time but I want to pay to to pay less. So you're going to have a 50% discount but your request can be can be delayed and and so on up to a few minutes. And on the other hand you have priority where you are going to pay twice the price but at the same time you're you're kind of guaranteed that it's going to be fast because you will you will have the fast track like in the airport or anyway anywhere else.
So yesterday I added that and I'm I'm going to use that to be certain that it works well for me but for you you might want to save a few bucks and not add that year. >> How much more expensive? >> Oh, twice. >> Twice, okay. >> Yeah. Um and I'm not sure it works with VO anyways, which is the most expensive ones of the model. And uh no with layer layer yet, so. >> Okay, yeah, yeah. >> Um so anyway, so I'm starting with I'm treating this this chart and what I'm going to do is to send it the first message that basically says uh I'm feeding you the whole book.
Uh I don't need to do anything with what with the book yet, but you have it, it's in your context and instructions we will follow. And then we're going to define a style. Uh usually I just set nothing and then I let uh Gemini come up with its own style. Uh but then but I'm getting tired of having exactly same size or same style always, so let's try something. Um uh a colorful um building block style. Um Well, let's [snorts] let's go with that.
Let's see what's how it goes. So I'm just defining a style. And by the way, if you're using uh Colab, this uh Colab has those things that they called um they called magics and that's kind of nice to to create those notebooks and to have the like forms that people can fill and that's just that fits into the into the code. Um then some system instructions uh to to to uh to direct the model uh like to into the kind of images that we want cuz the problem I had at the beginning when I was working on those examples is whenever you ask a portrait image and the model knows it's about a book, it tends to create like cover pages and to add the title and that's I didn't want those styles or to create ones with different panels and I didn't want that either.
So just that's just what the system instructions are about. And then like we can start working and basically I'm going to ask it to create prompts for each characters of the book. So can you describe the main characters only the adults? Actually you we could remove that because it's was just from the beginning of Nano Banana when you could not generate kids images in Europe, but it's not it's not true anymore. So um you can you can create image from nothing with kids, but you can't edit images with kids.
That's that's a current limitation in Europe. But anyway, here's all uh all prompts for for each character and then we can move to uh creating the the the images for for each of them. And what I'm going to do is that I'm going to create another chart just for the images cuz I don't want to mix the the text and the and the images output, but I want it to be a chart so that it will have all of the history of the previous images it it created.
Um so I'm I'm setting up with the responsibility image, the aspect ratio you want, the system instructions we decided did and the style and priority as as I said before. >> [snorts] >> Um and here we go. So here's the mole. It's nice. Here's the uh water rat. And that's actually way faster now that I'm using priority, so this works. Um I I could have done it in a better way and like make all of the calls asynchronous.
That would have wouldn't have worked with with chat mode, but that would have been a way to to make it faster. Um I think it's good enough. The the toad Ooh, way bigger than the car for some reason. >> It's the chatbot. >> Yeah. Uh and then Mr. Badger. Um Yeah. So, is he dark? This is mainly a demonstration of what you can do. Like every time I run it, I'm thinking like, "Oh, if I was to optimize it, there is plenty of better ways to do to do that." Um but still And the badger and then the otter will be at the end.
Um One of the thing I'm doing in the code, if I go up, is that I'm also saving all of the generated image from the character in a in an array, and I will explain afterwards why I'm doing that. And that's once again, I'm just appending them one after another. It In in real world, I would I would save them in a in a better way. Um So, and up and the otter is here. So, now we can do we can move to the next phase, which is basically illustrating the book.
So, same thing like I'm going to ask the the chat for each chapter, give me a prompt to illustrate the what's happening. It should be a single image, not multi-tiled one. I'm trying to force it to describe the character again even though there I will have the images as references just because it's it's it's still better. So, let's go with that. That should be quite fast. Um Well, same thing like the the the way it works with chat, it's uh, it's keeping an history that is basically all the previous messages and it's sending back the history to to the model every time we make a we make a call, um, which can be like in this case since it's it's basically resending the book all the time to the model.
That's that's the reason why it it can be quite slow. Um, we actually uh, released new API a few months ago that are called the interactions API. Um, let's run that while I talk. Um, and the and the main difference between the the new API and uh, and the old ones is that the new APIs are going to be stateful and stateless. And what what what it changes is that every time you you make a call, you get an interactions ID.
And you can reuse that interactions ID in future calls and it will recover all of the all of the context directly from the server. So, you don't you don't need to re-upload the same context again and again and again at every turn of the conversation. And it's also making it easier to um, to uh, to fork the discussion. Uh, for example, like you create uh, you want to create a song and with with images and uh, with a cover image, you can create the lyrics with one model and then we you you you you fork it and then on one end you create the images and the other end you you create the the song.
You had a question. >> Uh, how long is the order the session? >> Huh? >> How long is the order the session is it? >> Um, that's a good question. I think it's 2 days, something like that. Um, it's still it's still in in preview, so uh, but uh, there's good chances that at IO we make it the default API. Um, but I yeah, I don't I'm not using it enough to yeah, to know. Uh, one of the cool thing it does as well is that since uh, it knows that it's context that you're going to reuse, it's automatically caching it as well.
So, it makes it cheaper to to run. But, even though those the normal API are also doing the same, but um So, we have our chapter. So, the first chapter is next to the river. So, they the mole and the and the and the forgot his name characters are together. Then, it's on the road with the with the toad. And then in the forest in the snow. Woof. Scary. So, see and that's and and basically what I what I did there is that I used the fact that we were using the the history to to to trust the model to have all of the all of the previous images of the character so that it would remember how they look like and what's and how to uh create new images of them.
Um but there are actually better ways to do that. So, I tried another way, which is basically to create a new a new structured output that is basically the name of the chapter, the prompt, but also the list of the characters that are appearing in this chapter. And that's actually what I was doing if I was wanted to do it at scale with like more than a few characters. And basically I'm um I'm going to ask the model to to give me prompt for each chapter, but also to give me the to have thanks to the list of characters, I will only give it as references to the right images for the for the chapter.
So, I'm going to basically get the same thing here. Yep. And so, the first in the first image I should be the mole and the water rat in the third one in the second one Mr. Toad, mole, water rat, gray horse and and the third one mole and water rat again. Um so I created a very dirty uh script that basically uh searched through the the character images that we that we saved earlier so that I can give it the list of characters and it was going to to give me the list of images to to to give to the model.
Uh did I run it? Yeah. And and basically I'm going to do the same thing and go to and generate images for each chapter but this time instead of on relying on chat mode I'm going to use generate content so the like you know I recall but I'm going to pass it the images that are from the characters that are specifically in this uh in this chapter in this chapter image. So it should give slightly better context to the model of about what to um what to what to display and how to show them.
Um Let's see if it works better. Think I And like like I said like if I wanted to do it at scale I would I would have a lot of improvement I would do about that and one of them would be that I think I would generate more than one image per per character. I think I would have been one portrait image and then one full full full body image and then maybe from the side from the from the back and ask the model to tell me which exactly which like how they are going to be displayed so that I can give give the another banana exactly the the the reference we need for the for the for the other generation.
So, here they are. Um And two bits more or less the same, to be honest, but uh except that this time, I don't know why it's uh it seems to be attacking them all for some reason. >> [laughter] >> I don't think that's what the story is about. It's better uh you you may you might know better than me. >> [laughter] >> Um We can check the prompt. What What does the prompt uh say? Blah blah blah No, it doesn't say anything about attacking the No, it's rescuing him. >> [snorts] [laughter] >> Yeah.
Yeah, that's uh that's basically how you can uh you can use the um uh use the model. Uh for For those who came, the content I'm showing, you can you can you can open it there. And uh it's uh it's a collab that is about taking a book and creating uh images and videos to illustrate the book uh and the characters. So, we went through um creating prompts for each character, and then generating images for each of those characters, and then creating a new prompt for each chapter, and then uh creating uh images for the chapters uh using the reference that we have from the from the images.
If you can't see what I wrote, it's goo.gl/cookbook-illustration. Um So, uh and then we're going to try to move to the next phase with the with videos now. Uh so, we are going to use uh VO to generate uh videos of those uh of those images. So, I'm going to swap to the bigger model because I can pay for it. You can you can skip to the stay to uh stay with the cheapest one if you don't want to spend too much. Um and basically I'm going to just do exactly like take the last chapter, take the last image, and and basically send it the same prompt and the and the reference image.
So, that's what I do here. So, when I pass an image to VEO, that's basically use it's going to use it as the first frame for the video. Um And funnily enough, like most of the most of the like a lot of the training for all video generation model is actually image generation because I think the most important part of generating a video is generating the first frame so that it knows where to start with and then what to do with it.
Um So, >> So, what is the best video model? Is it >> It's the one that doesn't have light or fast. So, >> So, fast is just the faster version >> Yeah, that's that's that's >> Even more expensive but faster. >> Yeah. So, VEO 3.1 is is the main VEO model and the other ones are basically like smaller versions of it that are running slightly faster and like basically doing less less generation turns. Um yes, I said I was going to repeat the questions and I forgot to the questions was which one is faster the more the better model among the three.
Um I guess the door has just opened. Um It's very early. So, [snorts] up we can see where it goes. >> Oh, thank you, water rat. I thought I was >> Hey, it saved. And and there was sounds. I don't know if we Can we have more sounds so we can see what >> Quick, give me your hand. We must get out of this dreadful place. >> Thank you, Water Rat. I thought I was done for. >> That's That's not that bad except the wrong character is speaking.
That's the problem with basically using the same prompt when you when you create the videos and the and the audio and the and the image because it doesn't have the extra content about what exactly is meant to be happening afterwards. So, that's why I added another I added another example that is basically working to the same except we are going to add one more step that is basically asking Gemini to come up with a new um a new prompt just for the just for the video and to explain what's what's happening afterwards.
So, um that's what I did. I'm going to animate this the this this chapter image and can you create a prompt review about what's happening in the next few seconds after the initial image. Um and I'm passing it the image the last image that it knows exactly where to start. Um And while it runs since we have yet again new people who came, if you want to follow on your laptop, that's that's the link to open the collab I'm showing.
So, google.gle/cookbook-illustration. And what we're doing what we are doing at the moment is illustrating um the book that is called the Willows the Willow something like that. Um and uh and creating images and then our videos to illustrate what's happening in the book. So, it came up with uh with a prompt that is like in a colorful building block style a water rat in this uh round brown face and blue jersey lowers his silver pistols and offers a reassuring pat to the mole's shoulder.
The mole in his black velvet smoking suit exhales the puff of white plastic vapors in relief and as his pink snout twitches. They turn and begin to walk together, blah blah blah blah blah. What's What's interesting in the prompt is that I feel like it's it got the the the style of the book. So, the prompt is kind of written in the same style as the same oldish style of of of speaking as well. Um and we can see that it's realized that uh um it's um it's not attacking the mole, it's just saving it and then and so on.
And um and also what's what we can see from there is that I you remember I said I said some system instruction at the beginning saying that whenever it described characters, it should always try to describe like what they look like and all. So, it added what all the address all the time, so which also helps with the character consistency. So, let's see if this one is better. I think it's better except there's no there's no discussion.
So, you you can try with yourself. Just be careful as I said earlier, video generation can be expensive, so don't don't try it like a hundred times until you get it uh it can be it can be that that's something. Then, um we have this new uh Lyria model that we shipped 2 weeks ago that is our new music generation model. And basically, as I said, it's just a model to which you can give a prompt and it will come up with with a song.
And uh and you have two different models. You have the clip model that is creating 30-second musics, uh like I think it's four four four cents per song. And you have the full song model that is creating uh up to 3 minutes songs, and it costs like twice as much, so 8 cents per per music. For the sake of the demonstration here, I'm using the clip model, so the fastest one. So, we're going to do to use exactly the same the same trick as before.
So, we are going to have Gemini create the the prompts for the music generation. So, so I'm asking it to create instrumental songs for each chapter. And to create to create them for Lyra. Keep the consistency between the chapters, but at the same times highlighting the what's what's specific in each chapter, so that they are like I don't want like three times or four times the same the same song. I want different ones for for each song.
And and by the way, something that I forgot to say is that as you can imagine, like we have like multiple models, we have multiple teams inside of inside of DeepMind, but actually all of them are working together, and a big part of the training data for the Gen Media models is is being made with the help of Gemini. So, so basically all of our Gen Media models are trained with like prompts that are written by Gemini. So, that's also why Gemini is quite good at at creating the prompts for for the for the Gen Media models because yeah, it's they they they have been trained to listen to him very well, so that's that's why this kind of of tricks of having Gemini write the prompts for you works quite well.
And and in any case, deep down there's always a bit of rewriting of your prompt that are being done by the Gen Media models before it's actually starting to generate just because otherwise when people are sending like one-liners, they're they the models won't won't to going to do anything interesting with a with a one-liner. And um so usually the longer your prompt, the more interesting it's going to be, and the more likely it's going to be following what's your um what you're asking for.
So, I think I talked a lot because it's like music generation is actually quite very fast. So, we can see the different songs. >> [music] >> I think that fits with a pastoral suits uh with spring representing spring and flowing waters. >> [music] >> And see next one, the open road. Feels more adventurous, yes. And then the dark forest, let's see. >> [music] [music] >> Well, and you can see you can see the prompt here like it it comes with which uh which um instruments to use, how to use them.
Um what's uh what's interesting with the Lyria model is that it's actually you everything is managed in the prompt. You don't have that You don't You don't have parameters at all like at the moment. So, if you want the song to be a certain duration, you can just ask in the prompt. If you want the song to be using what a certain scale, you can ask it in the prompt. If you want uh certain BPM, you ask in in the prompt as well, and the model is really good at understanding everything you ask in the prompt.
And you can ask like doesn't make sense for 30-second songs uh much, but like if you were you're building uh longer song, you can you can say during the first 30 seconds, that's what I want the songs to be and then you switch to something else after or you can say this is the intro, this is the outro, this is the always forget how to say that in English but the the part that's that is coming up multiple times in the song so that it knows exactly how to chorus.
Um this knows that this part needs to be repeated and that's it needs to be the same thing and the same model. And if you want lyrics, you can either provide the lyrics or um or uh or just let it invent lyrics. So, let's say we are going to change that. We are going to create songs with chapter um add lyrics to describe what's happening up in the chapter. Up. And let's see how it goes with lyrics this time. As I said, music generation is quite fast like especially the 30 second model takes a few seconds to to generate things.
The the longer part here is actually sending sending the book again to the model and to ask it for um for prompts. Yeah, I think it didn't I think it didn't work. I think it didn't >> [singing] >> Mole flung down his brush and ran [music] to [singing] the sun, away from the cleaning and work to be done. >> [music and singing] >> And see this one. >> Toad [music] in his caravan [singing] of yellow and red dreamed [music] of the dusty [singing] high roads ahead. >> And like you can see that the the theme of the song is quite the same that's I kind of the the price we have to pay for using chat mode because it remembers the prompt it did before so we ask for new prompts but it's still has a memory of what it did before so I think it's kind of anchored it to kind of use the same kind of prompt and describe the scene the same way.
But still it's it's very funny to to work with the lyrics and you can as you can see in the prompt it's basically just like adding the lyrics in the prompt and the model understand that this these are the lyrics I need to I need to play and and add in the song. And you can also say that this this specific part this is is said at exactly this moment in the in the song this part is is is is being said at another moment and if you check the output of the Lyria model you actually have the all lyrics with the times so you can create karaoke app or something like that using what you get from out of the model.
Let's let's try the third one. >> Into the wildwood where the shadows are [singing] deep and evil faces through the hollow trees [singing] peep. Mole is in terror and lost in the snow till Ratty arrives with his pistols aglow. >> We could do we could do musical with that. So that's that's for music generation And then, um we have uh we also have a text uh generation models. And I'm going to show you something uh very fun.
I guess you all know about the all text-to-speech model because everybody loved the uh notebook LM integration that can create um uh podcast. And uh and that's actually great that you can select two different voices, so you have to have two two characters talking with each other. But, I'm going to show you a trick that uh with which you can actually create something that is basically uh you you can add more characters than actually two uh when you're creating uh discussions with the TTS model.
So, here's what I'm going to do. Uh I'm going to ask the model to extract a specific dialogue from the from the book just because I didn't want to copy-paste it. Um so, I told it that it starts with small, neat ears um and a thick, silky hair and ends with his ears in the air. Um but, and I asked it to write it as a play so that it's basically a transcript of what uh which which what which character should be should be saying.
And the and the trick is that I'm asking it to create a specific style of the of way of speaking for each character, even though it's going to be using the same voice. And to and to write the transcript that way. So, when it's a narrator, I'm going to use one specific voice. And when it's all of the other character, it's going to be the same voice for all of them. Um so, uh so, narrator is saying something and then character is saying something and then write the style between parentheses, and that's what is going to tell the model how to how to how to talk.
And that's that's something that you can also use to say, "Oh, he's he is saying this part very like whispering. And then, this part is like has a lot of emotion in it." And you can you can play with the way the character talks. But, I'm I'm going to uh use that to ask the model to come up with very different ways for from the same voice to to to talk when each character is talking and that's actually creates the feeling of actual different voices for for each of them.
So, and then I'm going to pass that to the TTS model and to and to ask it to to read it basically. One of the trick and like I got I I lost 15 minutes about because of that yesterday the evening so this you cannot just give it the text to read. You always have to start with read this text or something like that. Otherwise, it for some reason it doesn't know that it needs to read the text that is giving it given to it.
And this is a very complex. I think there's no simple way to to set it up, but basically what I say is speaker narrator is using the the voice Sulafat and character is using Fenrir. And this is all text. So, narrators talks and character the first character talk and that's is going to have long poetic pauses and then the second one is breathless and unbolstered. And you can see that we can we can guess which character is which one because it's the same way of speaking that is reused for each of those lines.
And it's still running. The problem with the TTS model is that it's a very good model, but it's not a very fast one. The reason for that it's a >> ears and thick silky hair. It was the water rat. Then the two animals stood and regarded each other cautiously. >> Hello, Mole. >> Hello, Rat. >> Would you like to come over? >> Oh, it's all very well to to talk. The Rat said nothing, but stooped and unfastened a rope and hold on it.
Then lightly stepped into a little boat, which the Mole had not observed. The Rat sculled smartly across and made fast. Then he held up his forepaw as the Mole stepped gingerly down. >> Lean on that. Now then, step lively. >> The Mole, to his surprise and rapture, found himself actually seated in the stern of a real boat. This has been a a wonderful day. Do you know I've never never been in a boat before in all my life. >> What?
Never been in a You never Well, I What have you been doing then? >> Is it >> [laughter] >> not so nice as all that? >> You can still doubt that that like you you you can you can see you you couldn't guess that we are using the same voice for for the two characters being like they clearly were steered into different directions and you can use this trick to actually uh create like multiple characters, multiple voices for those characters and uh and make it seems like seamless for for for users.
As I said uh earlier, this is meant to be just demonstration on how to do it. If you want to If you were to do to do it like uh at scale, that's exactly I would not do this kind of like create uh trick like that. I would actually create uh a full transcript with the actual names and then keep on the side a prompt for each character and maybe sometimes you still want to uh you you don't want to talk them exactly as same way because sometimes they still need to be excited even though they they talk very slow and and so on, but still that's uh that's uh that's just to show how good the the TTS model is at creating uh different voices.
And even though I asked it to force an accent, it didn't it didn't do it, but you can also play with like this character has an Irish accent, this one is English, this one speak with a German accent or whatever, and that's that's also a very easy way to to create different feeling about the about the character for with using the same voice. Um we're nearly at time for the for the questions, but like uh just to finish like uh we we use a very large context window of the model to to feed it a full book and to feed it multiple times a full book because we've chat we we just uploaded all the time, but since it's it's a multi-modal in model, you can you just it's it's work with all the things and just text, so you can just feed it like an audio book.
Uh you can feed it video as well. You got You can play with dots and not just get I get limited to to text to illustrate on things. So, this is another example with with another book which which one which is The Adventures of Chatterer the Red Squirrel. And and we are basically going to do the same thing. I'm going to run all of it at the same time. Uh and this time I basically ask it somewhere uh to use a style that is futuristic science fiction utopia saturated neon lights.
So, it's going to be not the kind of squirrel you are expecting. No. Um Yeah. And while it runs, I think uh oh I said I was going to show you um So, we also have like I think you all know about AI Studio, but uh in AI Studio we have a gallery with lots of uh example apps that we are building. And I wanted to show you uh it's going to be in Gen Media. >> [snorts] >> Uh as I told you, we have the Lyria model that is creating musics uh songs with but we also have the uh no not this one.
Let's let's go with the uh Up. This one is better. We also have the Lyria real-time um model I was talking about and and basically you are asking it to create uh to to make music that is post-punk tunes and neo-soul at the same time. But you can [music] say, "Okay, I want more K-pop." And slightly more [music] drums. And I don't know what post-punk is, so I don't want it. See, you can hear [music] the music changing.
And let's go with something like more chill. >> [music] >> So, I think >> [music] >> as I said, that's my favorite model because I think it's underused and I I I like there's plenty of things that I >> [music] >> I can imagine doing like I as I said, I come from the video game industry. So, one of the thing I would have tried [music] is can you create music in real time for the player depending on like where you in which region [music] they are in are they in the forest are they jumping are they cooking are they fighting uh how much HP [music] do they have and so so so the music could change in real life in real time.
And that's so that's [music] uh yeah. That's the kind of thing you can can try. And the for some reason the link is not there but there's another very cool example uh our colleagues who are working on this model made is basically uh you are in space and each planet is a prompt. And you can move through the planets and depending on which planet you are close to, the music changes. So, you can just move around the planets and sometimes there are weird things happening because like Christmas songs is just next to Viking metal.
So, the the mix can be can be quite funny. Um so, that's it for the for the presentation. I have some time for questions now. >> [applause] >> Yeah, and for those who arrived too late, like you can you can check the the content afterwards. And as I said, we have this this cookbook that is basically a GitHub repo where we are adding uh uh quick starts on how to use the models, uh kind of some some tricks, and also examples of like more complex more complex things you can build when you are mixing different capabilities and models. >> I have a question.
I I don't know if you can answer this, but >> Yeah, I think for questions we need mic so that it's I don't need to repeat them. >> One two, one two, one two, one two, one two. >> So, So, thank you first of all very much for this nice demo. It was a lot of fun to follow along. I have a question. In our company, we offering to all employees also some of models and I think we are still on Nano Banana one because we can only offer models hosted in the in Europe and basically all the new models are still in preview.
So, we don't have access to Do you know if this will change? >> So, [snorts] the short answer is no. >> Okay. >> Uh but I like I was expecting the question because I really like I it's a it's a pain for everyone in in Europe. As I As I said, my job is to bring the feedback from the developers and to try to make things change. So, that's that's part of the one of the fights I'm fighting at the moment so that we uh we have some some ways to offer better better exit for for or better ways to use a model for for people in Europe cuz in Europe we care about like data privacy and data sovereignty and all of that.
So, I know it's a it's a problem. Um so, the the the core of the problem in a way it's the rule that uh at Google Cloud that every preview model is only available on global endpoints. Um so, that's uh unlikely to change, but what we are going to try to change is to release uh the model in global accessibility uh faster. The The problem we have we had with Nano Banana 2 and the Pro and uh and like Gemini 3 3 as well is that we release models too quickly back to back and so instead of like having Gemini 3 going GA we released Gemini 3.1 and so we like kind of reset the the pre the the preview counter.
Uh so, with that we we need to we need to make to do something about that. But yeah, I I know I I hear you. It's it's kind of my P0 thing that I want to change. >> Thank you. >> And like have you have you run the notebook at the same time? >> Yes. >> Which style were you Oh, wow. >> [laughter] >> Did someone else run it at the same time and then with a different style and that or maybe a different book? >> I mean >> No. >> I I did for Frankenstein. >> Oh. >> And choose like read or game. >> Sure. >> So so I did the Frankenstein book and choose like a retro gaming style and it was quite interesting.
So um Yeah. So the character looked very video game-like. >> Oh, yeah. And um Yeah, the the main difficulty with uh with books like Frankenstein is that sometimes the model is not going to be uh willing to uh things that are too likely to say graphic that could be happening in the book. So uh it can be a bit toned down or worse, it's not going to uh to accept uh to show the image. Um yeah. That's why I settled with kids books for the for the example.
That's easier. Ex- except when I was not able to uh make uh images of of uh children's, which was also another limiting factor. Um You know, I can show you order I don't know what's uh page afterwards is going to show. I can show you on other cool demos that we have ready to Gen Media. Um in the meantime, if you have questions, just uh raise your hand and uh and we can up. Uh up. If we're going back here, see, that's uh whatever.
Uh that's the uh futuristic neon style version of uh of Chatter the the squirrel. Um I can show you um a bit more about Liyue because it's new. And uh you likely already know everything about Nano Banana. Um Up. So, as I said earlier, when you work with um with Lyria, the you always get uh two outputs. Well, if you if you set the modalities to be uh there, audio and text, you will get two outputs, and the first output is going to be the the lyrics, and the second output is going to be the music.
Um and that's actually one of the few model where it's very interesting to use uh streaming. So, when you do generate content here, you can use generate content under underscore stream. And what it does is that you receive the first part first, and then the second part afterwards. So, you get the lyrics first, so if you want to do something uh according to the lyrics, like creating an image or uh like giving it the song a title, then you can do it while the music is uh is still generating, and you don't have to wait for the full output to be to be there.
So, you get the you get the lyrics, and you get the timing. So, the this um this sentence is going to be said at the beginning, and then after 4.8 sec, it's going to say something else, and so on. So, you can you can hear >> is still [singing] and cold [music] up here. The mountain tops are sharp [singing] and clear. And then a streak of gentle gold. >> So, yeah. And then you can do And you can provide the same thing.
And here it's only like you can see maybe the last one is providing when it starts and when it ends because it wants to have like 1.2 second without uh without things said at the end. Um up. Uh you can also create images from music from images. So, that's one of the thing I forgot to do in uh in today's demo. I I should have also give the the the images from the from the chapter, so that it would have been used uh as a reference to create the crazy image.
So I use this picture of like grocery list for for making a pot-au-feu and then it will come up with a song about doing a pot-au-feu. What was the prompt again? [music] An epic song with opera voices about this quest. See, it's becoming a thing. >> [music] [singing] [music] [singing] [music and singing] >> And that's also nice that you can use multiple voices as well. I'm going to skip ahead a bit. >> [music] [music] [singing] >> The choice I like. >> [music] >> So but then and now I I talked about the interactions API earlier but being all new ways of using the API so that's an example using those.
So it basically was the same as the current API. So you give a model, you give an input and you get response modalities. Yeah, but uh Philip who worked on that is not there so I can say it. I would not have renamed uh content to input because it's going to confuse everybody but that's how it is. Um And then you get the output and uh it's um and if we can check in the output, I think no. Uh yeah, we don't we don't see it here.
But there's uh and then it works basically the same way, except you get this uh interactions ID that you can use to uh to chain things. Um and then that's uh yeah, that's the one with images and prompting. Um what you can do as well is you can use the BPM part, so you can you want a song that is very fast or very slow, so that's that gives you this. And I even uh told it to use a reset accelerando [music] illusion, so it gives the uh the illusion that the music is getting faster and faster and faster.
So, if you want uh music for when you do your sports routine, that's uh how you do it. And as I said, you can give like uh specific time, so the first 10 second are going to be fast acoustic guitar and then it goes into piano for 20 sec for 10 more sec and then it full band afterwards, so Actually, it's [music] not following it. Okay, forgot forgot what I just uh showed. For some reason. And then but the easiest way is like this.
You can use uh the you can give the structure, so that's how I want my intro, that's how I want my verse, that's what I want my outro. And that's 30 seconds song, so it's not that good. You won't have the chorus, but you can also add the chorus and uh the bridge and so on, so um >> The darkness [music and singing] breaks the shadows >> What is that about? Oh, yes, the song setting. >> Hear the >> [singing] >> dawn's triumphant [music] call, A golden light, >> [singing] >> a glorious sight, chasing [music and singing] night with heaven's might.
Our hearts resound >> Yeah, I should need to make me a little bit of like longer song, I think, for this example. And but you can also chain all of that into uh everything together. So, this is a full song where uh from from the uh first 2 seconds I have an intro I I can't tell exactly which uh which scale to use, how intense it's going to be, and then it move to another verse, and so on and so on. And that's uh how you can get [music] something very uh complex.
So, it start very very slow, and then if we move as the drums uh and bass started [music] and so on, so we should be in this part. It's still laid-back. And and starting to be Yeah, to add grooves [music] and so on, so if you if you want to create complex things, it's better to use a full song model because it's uh the the the short one is taking [music] shortcuts to actually build something that is interesting in 30 seconds.
And as as I already showed, you can provide the lyrics, so it's uh creating a song about nano bananas. >> Yellow peel, a tiny sweet, [music] the nano banana, a tropical treat. But wait, it hums, it starts to create, switching into AI mode. >> Um but uh and and what's what's funny as well is that you can use it to create things that are basically not not uh songs. So, you can uh you can ask it to create uh a music but without with very calm uh music in the background and just some something reading a text or or or something like that.
So, this And this one I'm using the reasoning capabilities of the model to and is it's knowledge about what Shakespeare is doing and so to to create a text that is basically something that looks like Shakespearean. >> [singing] [music and singing] [music] >> And I didn't really tell it to read, so that's why there's still music, but you can you can really steer it into not having background music at all and I'm just uh say things.
And you it works also in different languages, so you can uh you can just ask it to uh to uh to create songs in all languages that you want. Sometimes there are a few words that are not pronounced the right way. It's still uh it's it's still getting better. Uh but then and this one I tried to ask Rabbit to uh to use two [music] different languages in the same song and since it's the same as uh TTS and live models that we have, it's really good [music] at switching uh switching the the the language in the middle of the of the generation.
And once again, I'm using the model knowledge of things because it's basically [music] trying to explain how bubble sort is working uh in music. Um and instruments I told you you saw it. So, yeah, there's plenty of uh very cool things to uh to build with uh with the Lyria model, so give give it a try. It's uh it's very easy to use and because everything is in the prompt. So, yeah. Do you have any other question? No? What do you want uh page to show you afterwards?
She has She has 7 more minutes to prepare uh something new in your in your presentation. >> So, and also just for uh for clarification, I'm not sure if this was announced through uh to all of y'all, um but there are some electrical issues in the building. Um so, uh there uh y'all are like the lucky valiant few that made it here earlier this morning of the uh most of the attendees were not allowed into the building. Um and so, uh the uh we'll do the session that's coming up at 10:40, but we'll also be bringing back everybody in the afternoon who was going to be presenting in the morning to do kind of like a a whistle-stop tour of all of the Google DeepMind things for the afternoon workshop.
So, if you would prefer to come back for the afternoon workshop, you can. It'll just be at 1:00 p.m. >> So, and that's was the example I was talking about. I just I don't know where Christmas songs are, but uh >> How can you try this and buy a phone? >> It just search for space DJ [music] and it's available online. Can you move the put the songs uh higher? She wants this to to be nice. So, the model is not meant to do voices, but it can do some kind of like uh vocalizations like that.
Uh >> [music] [singing] >> This is not a robot. at using the same voice and switching the style of [music] music in there. And you have an auto pilot, so it just moves around and that creates music [singing] until you stop it. >> [music] >> Yeah, it's all on rock, I think. Nashville sound. >> [music] [music] >> Yeah, it's not moving fast enough. Let's move it. Let's move it a little bit faster. >> [music] >> Australian dub. >> [cough] >> Turn down the volume a little bit. >> [music] >> Australian dub, I guess.
No one knows Australian dub here? >> [clears throat] >> Um [music] yeah, that's uh that's a really cool model and the only thing is that the [clears throat] session ends after 10 minutes, so that you so you don't feel like running it at at the time, [music] but I can feel it should be good. Let's see. It's just using a button [singing and music] to speed it up. >> [screaming] >> So, I will stop with that. But yeah, give it give it a try.
It's a really cool model to play with. Um Okay. So, yeah, I guess that's it. I will still be around if you have other questions, things you want to discuss that you didn't want to be on camera. So, yeah. >> [applause and cheering] >> Thank you. >> [music]
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.