Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
Words
16,253
Runtime
1:47:33
Speaking pace
151wpm
Reading time
68min
151 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Hello everyone. It's just that Philip and I were both Germans so we thought it was funny maybe we we can do it in German. Actually looks like there's a there's a German crew there which is nice. Uh no no worries. We'll we'll we'll do it in English. We'll do it in a couple different languages maybe we'll find out. Do we have other languages in the room? Other nationalities? Yeah, what do we have?
76 words, the words spoken in the first 30 seconds at 151 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 1,367 |
| Average words per sentence | 11.9 |
| Longest sentence | 71 words |
| Questions asked | 200 |
| Sentences containing a number | 53 |
Most used terms
Filler phrases
1,108 in total: uh 266 · um 216 · like 191 · you know 124 · kind of 113 · sort of 68 · actually 43 · basically 39 · I mean 30 · right? 17 · literally 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
Hello everyone. It's just that Philip and I were both Germans so we thought it was funny maybe we we can do it in German. Actually looks like there's a there's a German crew there which is nice. Uh no no worries. We'll we'll we'll do it in English. We'll do it in a couple different languages maybe we'll find out. Do we have other languages in the room? Other nationalities? Yeah, what do we have? Shout it out. >> Italian. >> So Italian, Spanish.
Any Icelandic? No. Then okay, close close. Romanian, nice. Dutch, any Netherlands? Okay. Any which? Yeah, Hindu okay. Canada? No, okay. Bangalore? No. Uh Farsi, all right, nice. Czech, yeah. Brilliant. Okay, this is uh fantastic. Well, thanks everyone for making your way over here with with all your languages. Really appreciate it. We can we can really put the model to the test today which is great. Um yeah, hi. I'm Thor or Torsten for the German speakers amongst you.
Uh >> Hi, I'm Philip. Um only Philip. So before German and English speakers so >> Yeah, it's nice. We So we work on the developer experience at Google DeepMind uh broadly covering kind of Gemini API and also uh working in Google AI Studio as a tool for you know developers to try out the models quickly uh and then also the the API interfaces. To use the API, you do need an API key. Uh actually, who has an API key already?
Gemini API? Okay, a couple folks. Uh who has used AI Studio before? Google AI Studio? Okay, a couple folks, grace. Um maybe we can quickly Sorry? >> Anti-gravity. >> Yeah, that's good. Uh you actually don't need an API key for anti-gravity, but um So, if you don't have an API key yes, if you have your machine on you, I hope you do. Uh it's it's it's hands-on workshop. I do apologize they took away your tables. Um so, it's a very laptop situation.
Uh like literally uh laptop. Um I hope you didn't bring your Mac mini or you know, whatever. Uh but so yes, if you just go to ai.dev um And I mean there's you can also go to ai.studio, you can go to ai.studio.com, but we paid a lot of money for ai.dev, so please use that. Uh it just it just redirects to ai.studio, but um >> Let me >> So, ai.dev/API key? >> Uh yes, or on the corner left side, there's a get API key and then API {dash} keys.
This way you can >> Schlüssel. >> Uh yeah, it's it's my personal account and we cannot change language, so Uh you need only a Google account, so no credit card, no nothing. All of the things we are going to do is part of the free tier. So, if you go to AI Studio or ai.dev and there's some like sign-in form, you can like just use your Gmail, no worries, we will not charge you. Uh and then on API {dot} keys, you find at the top right corner normally something called API key create or create API key.
Uh and when you do it for the first time, you uh might need So, similar for me, uh I can import my projects or I can create my projects. I can just click create, give it any name. I mean, we can call it uh i e. i e. workshop. Uh and then create project. We have to uh translate it to English. Takes a few seconds, and then you should be able to create key. And this key will be used for the demos or the hands-on things we are going to do use later. >> Wait, it shows the API key. >> Yeah, I will delete it, so don't don't copy it.
But you can like >> I could be very fast. >> Yeah. >> Thank you for >> You can use that key. You can like uh put it into your bash RC or zshrc file, or later you can like directly inline it whatever you prefer. If you have one, you you can use this one. And since we have a few more minutes time, we really want to make sure everyone who wants to follow along, take your time, create your API key. If there's an error appearing or something else is not working, feel free to raise your hand.
Tor will come to you and will help you. And >> And yeah, just a reminder, it is a secret API key. Don't show Don't do what Philip is doing there uh >> That's why I delete my key again. >> Yes. So, we do get that a lot that people um leak their API keys. Mostly it's clock code that is like pushing the API keys to to to get up. So, I recommend you don't do that. No, just just kidding. Just remember, it is a secret API key.
Um so, treat it like a secret. Uh don't share it with your neighbor. >> Yes. >> Um but yeah, do create that API key now, and we'll give you a couple minutes to do that. And while you do that, or you know, once you finish creating your API key, I'd love for you, you know, to just briefly introduce yourself if you want to and just sort of, you know, let us know what you would like to get out of the workshop. Maybe there's a specific use case you're working on.
Um yeah, we'd love to kind of get to know you all a little bit as well. So, while you create your API keys, uh if you want to, feel free to just, you know, shout out uh your name or nickname uh and sort of what you're working on, what you would like to get out of the workshop. >> Yeah, we're from Rebel World, from the Thursday I podcast, working for Corvus racing biases and uh big fan of Gemma and yeah, Gemma 4 and Gemini.
I'm one of the few who are using it for coding and stuff as well. I actually use and I'm also running it on the classes actually. >> Oh, nice. >> Um and I want to see more of this, you know. >> Okay. >> And I want Google glasses. >> Okay. Yeah, we'll we'll see if we can get to the glasses by the end of the workshop. Sorry. Yeah. >> If you have a seat to your right, can you jump in so we can fill up on the side here? Makes it easier than people have to find one. >> And while we are waiting, uh we launched something cool on Chrome, which I haven't activated it, but when you go on your tab bar, click right, you can now move tabs to the side and we have vertical tabs now, so yeah.
So, more screen, that's good. >> Nice. Okay, cool. So, glasses. Uh are you are you are you running what are you running on the glass? Well, technically on the phone, right? >> It was the software we had. You used Gemini, which called Open Claw. >> Nice. And then that's just using Gemini Life kind of through the web socket on the phone, right? Yeah, nice. Yeah, that's cool. Yeah, Vision Claw, if you haven't heard of that, uh pretty pretty fun open source project.
Um and I think you can run it, yeah, sort of on the Well, now that uh Meta has opened up the SDK for the the Meta Ray-Ban, uh you can actually hook in something like Gemini Life API uh into the glasses. Well, the glasses connect to your phone and then the phone actually connects to Gemini Life, which is cool. >> but uh it shows what will be possible. >> Yeah, nice. Any problems creating API keys? Everyone has an API key?
Should we check? Okay. Uh anyone else? Anything specific that Yeah? >> Uh Michael, work for Digital Guide in Zurich. Um we're building like our first kind of AI customer support agent, so >> Cool. >> just interested in thinking the general how you guys are doing it, comparing. >> Okay. Cool. >> Nice. >> Last chance, anyone else? Okay. >> All right. So, we can start. Before we go into the hands-on session, we I have like 10 15 minutes slides um to give a bit of a background what we are going to use to build.
The first session, which I'm going to do, is more on like the building an agent without any live what audio input, that's where Thor is going to take over later, and then we are going to building some very nice conversational agents. So, um who of you has used the Gemini API to make an API call to Gemini before? That's a few. Have any one of you used the Interactions API? One, okay. Two, okay. At least some persons. So, the Interactions API is a new API we launched in December in beta, which hopefully will succeed Generate Content soon.
It's a unified API to use with models uh with agents, and it's much more aligned with, I would say, the industry. So, much closer to what you are familiar with from open models with the Chat Completions API or from Anthropic or from OpenAI. Um and what we are going to do the slides, then we build a coding agent like a small little cloud code with reading running reading files, writing files, running bash commands, and then we have some time at the end if you have questions around it and then we do a short break, toilet, drinks, and then Tor will continue.
We all have an API key. So, as I said, we want to build an API which works for both models and agents. And when we launch the interactions API, we also launched deep research. So, maybe you have used deep research in chat GPT or in the Gemini app where you basically start off a query and then you get back a plan and then the model goes on and like does deep research for 10-15 minutes visiting hundreds of sites. And the API supports both models and agents and it's very simple to switch between those.
You basically either define a model which could be Gemini pre-flash or you define your agents which could be deep research. And we are working on it to for you to bring your own agent or you to define your own agent that you can customize all of these behaviors. And it's the same surface. So, we'll see later, but when you send a request to Nano Banana to generate an image, you can like basically chain those interactions to a flash model to do something else to Luria to have you like generate audio and or even to hopefully soon video to generate you video.
And the interface is very similar to what you see from OpenAI. So, we now have like those content blocks basically. So, every input you provide and output you provide the same type. It has a type field which could be a function call, a thought signature, a text, audio, video, image which hopefully makes it for you easier to build with the Gemini API. And it is less, I would say, Google branded, less proto specific, less GRPC to make it easier for developers to build.
Um the core primitives of the interactions API, in addition to making it easier, we also introduced state on the server which will be which we will use for building our agents, so we don't need to manage our loop and always send back the whole history. Very similar to responses API, you now have a previous interaction ID you can provide, which basically attaches to the existing history, so you can like just send a new input.
Um as mentioned, we have an agent with deep research and background through, so you can start your research, can poll it, or um soon use web hooks to get notified when your research is done, so you don't need to keep the connection open. We have the type blocks, and then also we have the same like streaming pattern, very typical for web development using SSE, which also makes it hopefully easier for you to build. We all support the built-in tools, but we also now have support for remote MCP, and I think 2 weeks ago we launched tool combination, so you can now combine Google search with your own custom function, which was one of the big features people were asking for years, I think.
And now you can do this. Um And to summarize again the differences between So, generate content is what we have today. Interactions API is what we will have tomorrow. We will have um server-side state management, but you don't need to use it. So, if you say, "Hey, I want to manage my turns. I want to manage my context. I don't trust you, or I need to do like context engineering. I will need to remove certain parts." You can do this.
You send can like way easier to send new input, basically. We have the built-in agents. We have also background support, so you will get asynchronous execution. We all see with agents, when you send a prompt, it might take 1 2 3 4 minutes to complete, and keeping HTTP requests or connections open for, I would say, more than like 10 seconds is not a very good practice, so you want to like use asynchronous calls to either get notified or like do polling when it is done.
And it is less proto-oriented, which I really prefer. Um it's much closer to what people know from like the the developer ecosystem. And also, a side effect of the state management is that the implicit caching for the API is much better. So, for the one of you who don't know what is implicit caching, so when you send a request to the model, the model needs to encode all of your input tokens. And you can cache those encodings for follow-up requests to save cost.
And cache requests I think are 90% cheaper for the input tokens. And when you have to manage your state or context yourself, you maybe strip out uh line breaks, remove um certain parts, this breaks the cache. And using the server-side state, the server keeps the context, so the chances for your cache hit rate is much higher. And we see like two to three times better cache rates from like the the startups using Interactions API today.
Quick code example on what we mean with making it simpler. So, on the left side you have like the very proto-specific one-off input parts with inline data or text where the the field basically describes the type. And on the right side, almost looks similar to other APIs I would say, so should be much easier if you decide after the workshop to give it a try. Um then I guess all of you know roughly what an agent is. We have a brain or our model which decides what it wants to do, calling a tool, generating text, doing something else.
We have tools which basically gives our brain hands and eyes to interact with the environment where it is in. We have the context, it's basically all of the all the model knows uh what it has to do, what it can do, if there are certain preferences, if there are certain constraints. And then we have the loop which basically combines our model with hands and tools and runs it until the model no longer calls the tools and generates a text.
So, some quick examples on how to use the API and then we go into like the the nice hands-on part. So, basic chat usage with server-side state becomes very easy because we have our Interactions create call. So, we define our model, we define our input, What's the capital of France? We get back the output, and then we can just continue providing the previous ID. And the model behind the scenes, or like the server behind the scenes, basically has our user input, our model output, and then appends the new user input that you don't need to have like a client-side history object where you append the user turns, the model turns, and then the user turn again.
This becomes very helpful when you build agents where you have to loop and always need to append user input. And as mentioned before, that also works for agents and models. So, we really want to build this unified interface where you can continue your conversations, no matter what model you used. And in this example, we basically run a deep research request on research AI agents in 2026, and then we take the research output, and just continue with the Nano Banana model to generate a visual for it, and it's like four lines of code, basically, without for you to what's the context, how do I provide the input, and hopefully makes it a lot easier.
But as said, you don't have to do it. So, like the input field also accepts the same array with role user, role model, and all of the inputs. Um for tool use, it's also hopefully now much easier. We have a type function, our name for our function which we want to use, describe, and then we have the parameters the model needs to generate. And then here's roughly what we are going to build. So, we make our API call with the tools, and then we check what the the output of the interaction is.
So, requires action basically means you as a client, you need to do something. So, the output generated some function call or some object which you need to react to. We iterate over our output types, check for the function call, execute the function, append the result, send a new interaction until the model no longer decides it wants to call a function and generates a text or something else. Okay. Um The last time I did a workshop, we were all still coding manually.
So, that's my first time doing a hands-on workshop where at least I don't code much manually anymore. I'm not sure how you do it. So, what we are going to do today is we don't code manually. We are going to use your preferred IDE, agent, CLI of choice. I'm not sure how many of you are following, but um to make it easier, we created agent skills for our agents to use. So, if you go to the Gemini, you can search for Gemini API docs coding agents.
Ah, yes. Sorry, of course. Um I can also Can you see it? >> Yeah, I think you can also just Google it. >> Yeah. I mean, I can Gemini agent skills. Okay. Um Yes, this one it should be the set up your coding agent with Gemini MCP and skills, I think. That should bring us here. And I Really? Yes. Or if not, you can go to the documentation and then on the left side in the getting started section, there's a coding agent setup.
Are you Are we successful? Okay. Um >> Just one question. Should install this skill globally or just for the project? >> You can You don't need to install it globally. And that's a good question. So, how do you install? So, we have multiple skills. So, we have the Gemini API dev skill. We are not going to use this one. That's for the generate content API. If you want to use that API or if you are familiar with it, go for it.
But, we have the the Gemini life API. Um that's what Tor is going to use later. So, we also don't install this. But, then we have the Gemini interactions API. And here you can either pick the the first command or the second command depending on what you want. And then like just copy it uh open your workspace um where you are working in. And then I already did it, but you can like add the command NPX install. And then you should get a wizard asking you to install it.
There are many pre-selected. Don't get confused by it. It just means that all of those agents uh are compatible with the dot agents folder. So, you might have a dot agents skills {slash} Gemini interactions API. And in there we have our skill. Yeah. I let me also uh with So, I tested with cursor and an anti-gravity and also Gemini CLI. So, if you are using one of those three, you are good. If you are using Cloud Code, I think it should work as well.
I'm not sure if they follow dot agents best practice. If not, I mean I'm sure you can Mhm. Yeah, the skills, but I think they only look at dot cloud. >> Uh I think They don't look at agents. But, I think it installs into skills now. >> I don't know. Okay. So, >> I'm sure it will. Yeah, if you install NPX skill with cloud as the agent >> Okay, yeah. Question. >> Um you showed two different commands. One says skills at this age and the other one context seven or >> Yeah, they both both work the same. >> What but what's the what is contact seven then? >> Uh so contact seven is I think a product based out of upstairs.
So contact seven is has an MCP server and our skills CLI which you can use to get access to like skills. It's like a public repository. So and skills.sh is the same from the result. >> Okay. >> So it's like and it it works both works on GitHub. So the Google - Gemini Gemini skills is a GitHub repository. And the GitHub repository includes all of the skills we we documented. So you can also go there and find it there.
There are also the deals and installs commands and it makes it much easier than cloning it and making sure you have it in the right directory. And then what we can do is um to make sure our skill works, we can just like ask our agent uh what skills can you use? And then it should If it if you have installed it correctly, you should see yes we have our specialized interactions API skill. So anti-gravity here picked up our skill in the {dot} agents {slash} skill folder.
Okay. How are we doing? >> Can you zoom in just a bit? >> Zoom in more? >> One more, yeah. I think just in the back. >> Okay. Better? Yeah. >> Um I'm I was using Telegram with my apps and as a result of this being run on the system, so would I have to have it run >> No, you can like you can should be able to do it on like your open clock kind of stuff. If you tell it install the interactions API and then >> Just the final product of the workshop, will it run? >> Yeah, it should run.
It's a Python script we are going to execute later. >> So because I'm only running the box >> that's should work. Okay. Any any difficulties installing the skill? Any questions what a skill is? Everyone familiar with skills? Okay. What we can do maybe to give you some insights on what skill we created or what it contains, right? The importance when creating skills is it should be either something the model cannot do reliably or if you have some personal preferences on like how to do a certain workflow or I don't know, you always need to run tests using bun or something like this.
And what we did with our skill is we made sure that the agent is aware of which Gemini models are available. A common issue we saw before that is like Gemini always used Gemini 1.5, which is no longer the latest model. We also included the agents here. We have some like very high-level information on how it works, but we did not include like all of the documentation. What we did instead is you should see Yes, a link to our documentation, which is available as markdown.
So, instead of the need to always update our skill with like for example, we added a new feature to interactions API to combine tools. We would have needed to update our skill and then everyone of you also need to update your skill to be able to use it and then we are not making a lot of progress in terms of like knowledge cut off. Instead, we provide the information as part of the skill. So, all of the agents now have like web fetch tools, so they can query the information based on the skill and then like we only need to maintain like the documentation, which is mostly up-to-date.
Yeah. >> How do you find that workflow like efficiencies? >> Sorry? >> How do you find that workflow like tool calling efficiencies and all that sort of thing? Do you find it like it having to go and fetch the page to that >> I mean, it it's normally you would provide it as a reference on your local file, right? So, it either needs to do a read file call or a web fetch call, which is the same I would say in terms of like cost.
And it works very well. Okay. So, what we are going to do as a first example, since we are not white coding, uh we want to build something more substantial. We not just say build an agent, we want to be more specific. So, we want to build um create an agent class with a constructor and a run method. The constructor creates uh GenAI client. So, the GenAI client is what we are going to use to call our model. We also need uh the files uh model.
And we also need uh global previous interaction ID and then at the main method to run an example. Uh do all of that in uh workshop. Okay. So, then as we can see, the model in this case or the agent as a first step, read our skill, analyzed our main file, is implementing the skill. Okay, checks. It should probably fail. Yes, because I'm using UV. So, I stop and I tell it use UV uh from the workspace. And and then we let it generate UV.
Okay, it checks if we have installed a library. We have. So, maybe in your case, the agent still tries to install Google GenAI. If not, you can do it yourself with like pip install or UE pip install um google-genai and then make sure to Okay. So, we have our starting agent class. We have our GenAI client. We have our model ID. It defaults to Gemini free flash. It would never have defaulted to Gemini free flash if it would uh it it hadn't read the skill, right?
Because then we were stuck with Gemini 1.5. We have our run method which calls uh makes our interactions.create uh call. We have the input text. We have our um previous interaction ID. We set our new previous interaction ID and then we return the text and then nice as a main example. Um which we will run in a bit. We create our agent. We have turn one. So, my name is Phil. And then the agent uh runs it. And then what is my name to Now, we can check if our Gemini created Gemini agent uh works with uh multi-turn and our interactions API.
You might see a warning similar to the one I got here. Uh interaction usage is experimental. That's as we are still in beta. We really work hard to get the API out of beta to make sure you can use it in production. And our call was successful. Hi, Phil. Nice to meet you. Okay, maybe we weren't successful. I don't actually have a name. Also, what is my name? Hm. Maybe we should check. Did we do it correctly? I mean, my name is maybe Ah, okay.
That's still the call. Sorry. I asked it, "What's your name?" I mean, it's a language model for trained by Google, it makes sense. And what is my name? Your name is Philip. How can I help you? So, yeah. >> You're using Gemini 3 Flash as a coding model? >> Yes. >> Respect. >> I mean, it it it works really well if you are providing good instructions with good skills and good context and don't expect it to, I don't know, like cure cancer.
Like, if you have a very good understanding of what you are trying to build, Gemini 3 Flash is like very fast. I mean, it didn't take much longer. Didn't consume credits, right? Every one of us is somewhat token constrained at the moment. And it works really well. Okay. >> Use it for a generic use? Yes, but for planning and coding >> Yeah, no, it it works. I mean, we will going to use free well. I mean, like it also asks like if it wants to run, so we are closing the loop even like with Gemini 3 Flash, it tries to run our script to make sure it works.
And then we will continue in a bit. How are we doing? Anyone making successful calls? Yes? Okay, perfect. >> Back row. >> Okay. Any any issues? Any errors? Any questions? Yeah. >> I'm using TypeScript. >> Okay. It it works? >> Yeah. >> Nice. Great. >> Not so much. >> Mhm. Awesome. Even coding on a phone, so Okay. So, normally the next step for our agent, right? We now have our like very basic run, we can chat with the model.
Now we need to add tools to it. And we want to build some kind of a coding agent, so our first tools are we are going to add is a read and write file tool, and we just continue in our main agent thread. It's like, okay. Uh next, we need to add a uh read file and a write file tool. Create the create a basic Python implementation. And also the JSON schema definition. So, when you use function calling or tool use, right?
We need to create a JSON schema which we provide to the model. So, the model understands what it needs to generate. Once it generates that schema, we also need to have some kind of a code implementation which we can then run on the client. So, we ask it to create a Python implementation and also the JSON schema. Uh and uh map to for the key and key function comma schema. Okay. Let's see what it will come up with. Okay.
It's cheating a little bit. So, I have like a solution folder and it looked up the implementation implementation. Yeah. I mean, that's why it's still important to like check your work. I can like it it it found like the solution and then it was like, "Oh, I got the solid Python example in solution.h to guide me. The task is implement read file and write file." >> That's a smart model for sure. >> Yes. Uh so, what we got back is we have two new like very basic very very basic file implementation.
So, we have a read file tool with a file path which uses um Python syntax to open it, to read it, and we have a write file tool with file path and the content and writes it. And then we have our read file schema. Um reads a file and returns the content. Write file, writes the file and returns the content. And it made some updates to our agent. So, what did it change? Okay, our input is now a text uh string and a list.
Makes sense since we now need to return uh function call. We check. What do we do? Okay, we create a tool definition for our model, which is the T schema. So, that's our tools map schema. And then we have our loop, which you might be familiar with from the slide I showed. So, after we run our request, we check the interactions. outputs. So, the interactions.outputs include all of the events generated by the model. And since Gemini is a reasoning model, it also includes, for example, the thoughts and the thought signatures, which we need to return.
Since we are using the previous interaction ID in the server-side state, that's done by us. And we only need to check, okay, do we have a function call? We have a nice debug, so we can check it. And for our output so our function call, we check our tools. Do we have our tool or not? I mean, maybe the model wants to edit the file, but we don't have a edit the file tool. We would catch it here. It calls our model. And then it creates the tool results.
So, for function call, we have a function result. That's also part of the change. We really want to make it easy. Um, we will see later when we use Google Search, there will be a Google Search call and a Google Search result to have very the the same schema. And then we use recursion. So, if we have tool results, we basically call ourself the self-run method again. If we don't have tool results, we return the interaction text.
And it also updated our example, write hello from the agent to a file named hello.text, and return it back. So, we I mean the changes roughly look good to me. We can try and run it, so you will run Python workshop main. What do we get? We get our tool call with write file. We get our tool result. We get a read file call again and then the final agent response. The file hello agent text was successfully created with the content.
Okay. Agents are not only like singleton, right? Maybe we wanted to respond. So, next step is basically very ambitious, telling it I want to have uh continuous I don't know. stood in implementation to test. Let's see. It Let's see where it will continue. No, no, I hope at least it should update our main function where we have an input, a while loop, basically waiting for the input using our agents. So, yeah. So, we have our user input.
Uh and then we have we always continue with our agent run method, which then inside the agent runs it inner loops until they are no tool calls anymore. And if there's a result, we basically get back our response. Okay. Let's quickly read. Is it too fast or are we still on track roughly? >> It's fast. >> It's fast? >> Yeah. >> Okay. I mean, we will share the code later. Um also very happy to answer questions. Uh we have many like blog posts and examples of that online, so you can if you are interested in like rebuilding it later with less fast, um but I'll I try to be a little bit slower, but I also want to give Tor enough time to have you speak.
So, what we can do now since we are like our agent implemented our like while loop to um to to provide input, we can say something like hello, right? Normally, we should now not do any tool calls because hello is like a our agent correctly kind of understood or Gemini in this case understood, hey, hello is nothing I need to solve with a read or write file to write. So, I can say, "Hello, how can I help you? Can you write or maybe can you create uh CSV with uh thumbs up?" Might take a while bit.
So, it does the thinking and it does the function call and whoop. Okay. Um certainly, here's a CSV file, blah blah blah. Um okay. Can you write it to disk? Uh Yes, I can if you like. Wait, maybe our model is not having our tools. Let's check. There's the tools map. It does. Why is it not using our tools? What tools can you use? Okay, it has the read file. Yes, maybe we are we are not explicit enough. And what we can do to improve this in 1 second is we can add a system instruction to tell the model, "Hey, you are a coding agent.
You can use tools to write and interact with the Okay, there we got our tool call right file with our CSV. Cool. Since we saw our mistake, we cannot tell the model, "Hey, add system instructions for the interactions API call and add an sample prompt for a coding agent." Boom. And what's really nice now, since we loaded the interactions API skill in the beginning, the model still has the awareness of, "Okay, how do I add the system instructions to the interactions API?" And what I can guarantee you is that the Gemini 3 flash has never seen any code of the interactions API because the model was trained be- before we even released the API.
So, all of the work we were doing so far is based on like the skills and like the the coding infrastructure and been part of the the training. Okay, so, what did we get? Um okay, we can provide it on the run command or we have one when we create it. That's good. And okay, coding persona, you are an expert software engineer and helpful coding assistant. You have access to the local file system. Okay. Let's accept it. Let's start our agent again.
Let's say, "Hello." Okay, hello. How can I help you with your software engineering and coding tasks? I mean, definitely better than what we had so far. And what did we send before? We said, "Can you create SVG with a thumbs up?" So, can you create an SVG with a thumbs up? And now let's see if it calls yes. And now at this time we got our right file tool call. And then also we got a hey, I have created a thumbs up SVG file with a simple line out thumbs icon.
So, um can I Yes, there we go. We have got a thumbs up icon. Cool. Um and of course what's missing for coding agent right, we need to get some bash tools. That's now not part of the solutions folder. So, let's see how we will get our bash tool. Now at a similar run command tool that allows the model to execute bash commands. Okay. Creates an implementation plan. I feel like I think there's a question. Yeah. Okay. So, we have our run command which uses a sub process which in this case I guess it's okay.
We don't care too much about security for this example. And then our output is a stood out. Um works. Edit. We have our run command tool. Yeah, it even updated our system prompt and now let's stop that. Let's clear that. Run it again. Um any suggestion on what we should test? >> Is not a banana to create an image? >> I'm not sure that will work because we don't have any skills or any information for it. Time? >> Time, yeah, I guess. >> Um. >> From bash. >> Get the time.
Tool call run command date. Wednesday, April 8th. It looks good. Um, yeah, cool. That's our small little coding agent. >> Delete all files. >> No, I mean, let's not do this. Any more questions, any more ideas? We have like roughly 5 to 7 minutes. Yes. >> Just a quick question. So, because the state is kept in the side bar. >> Yeah. >> Is it possible to like fork the history? >> Yes. Yes, so Um. What you always can do.
So, what we are doing here is right, we always use the previous interaction ID from the previous turn. So, we basically stack it and you can always go back to any index in the stack and branch from there. So, if you would keep the interaction IDs on your client side, you can always use those to I don't know, branch out and like, I don't have like a first um, prompt and like do basic web search and then like use this as like a base for like five parallel requests doing some other work and you can always get the context.
So, we have an interactions.get method which you can use to retrieve the interaction and then also get the previous interaction ID so you can basically go back until the beginning and get all of your state if you want to save it for later and the default for those interactions being stored on the server for free tiers one day, for paid usage is 55 days at the moment. Yeah, you had a question. Sorry? Okay, perfect. More questions, yeah?
Uh, no. So, once the Gemini models have a million tokens context, what would happen now if you reach that, you will get an error. But, we are working on context compaction techniques, but it's easier said as done, and still something you currently need to maintain on on your client side. No, no. So, when you send the request, you get an ID, the interaction ID. And the interaction ID stores your input and the output of the model.
In the free tier, the input and output and the ID is stored for 1 day. So, meaning if you send the request now, and you continue 8 hours later from that point, the state or like the context is still available. If you would send a request tomorrow, it would basically say, "Cannot find request with the old interaction ID because it basically is pruned after a day." But, if you use a paid API key, it's stored for 55 days, and the interactions API is also coming to Vertex, and I think there might be a little bit more flexibility in terms of um how long you want to store or customize it.
Yes? We don't have one yet, but hopefully soon. Yes? Sorry? And we can also use Gemini, which has a TTS model, which can speak the But, speaking and listening, I mean, tool will show many many cool things. Any questions regarding the interactions API in this small little agent? No? Okay, then you get 8 minutes. >> It's a 5-minute >> Okay, 5 minutes. >> bio break and then we'll be back here >> with >> to make your agent talk. >> Yes.
Cool. >> Cool. Thanks. >> That worked really well. >> Nice. Yeah. >> I used the old one and I'm super happy that now there's a new >> Mhm. >> It was 2.5 before. >> Yes. Yeah, it's it's been a while. Yeah. How much better? >> Yeah, much better. >> Much There we go. We didn't even pay him for it. There you go. >> Big upgrade. >> Big upgrade, yes. Um Yeah, maybe uh I know there's a couple more minutes, but >> Can I Can I ask a question while we're waiting? >> Caching? >> Yes.
Yeah. >> How is that Are those cached at the context array element level or the whole context Like when adding the elements to the interaction Is each one of those individually cached? >> That's a good question. We probably need to find Philip to answer that. I actually I actually don't know. Philip, caching question. >> Yep. >> Uh the input tokens. >> Yep. >> What was it? >> Um context cached at the individual interaction level or is the whole context >> It is not on an interaction level.
It's more on an like object level. >> For example, when you provide an input like no interaction ID, first input PDF 4,000 tokens and the text input with 10 tokens and you do an follow-up interaction call. Maybe only the PDF will be cached and not the other like the short text. And then if you do another one, maybe the PDF and like the follow-up turns will be cached. So it's like more on like an object level. But since you I mean how is it It's very easy to make a mistake in caching if you like even the slightest change in your prompt removing white space, line breaks will break it.
So like having this rely on the server to keep it, it's more guaranteed that it's secure and it could be as easy as hey, my user says there's an empty line break at the end. Okay, I remove it and then I use that history again and then it falls apart. Yes. Uh optimally, but I mean it depends on where your request kind of hits it and like how fast you follow up, but the cache rate should be pretty high. >> Cool. Maybe to kick things off, we can sort of look at one of the examples that that we can build uh with the the life API the new model.
So this is now Gemini 3.1 flash life which came out 2 weeks ago, I think. Um very exciting. It's been a while. I think the previous 2.5 native audio model was December. So it's been a while. The reason being we did kind of major underlying architecture. We worked to you know, ideally lower latency, you know, better scalability in the in the future as this ramps up. So that's very exciting and then, you know, in parallel we're working on bringing the life API to the interactions API as well.
Um so that's something that Philip and I are working on at the moment. Not there yet, so we still have to use the uh Life API for now, but let's do a little demo. And so this demo uses the Life API in combination with Lyra 3. Uh so on the Gemini API, have have people played with Lyra? So the new Lyra model is actually a music generation model. And so Lyra 3 can now generate songs, uh you know, including songs with the lyrics.
Uh so this is the clip model, which is kind of a 30-second uh clip generation. Uh but then also we have a full song model now, where you can generate kind of full full songs. So the idea is that we have our conversational agent, which is kind of the DJ that you're interacting with, and then that DJ has a tool call to generate music, right? Um And so, you know, I So I'm originally German, but, you know, growing up I learned kind of English through Harry Potter and baby BBC Radio 1.
Uh and if you remember back in the day you could like phone in and like wish for a song. And so that's sort of the the the idea of the experience here. >> Right then, you're on live jukebox. What absolute banger are we cooking up today? Tell me, what genre are we leaning into, and what sort of vibe are you going for? Let's make >> Can we get some sort of high-energy German um techno schlager about the AI scene in the UK? >> German techno schlager about UK AI?
That's absolutely brilliant, mate. Proper Radio 1 style. Any specific lyrics or themes you want crammed in there? >> I'll just surprise me. Let's go. >> Right, you beauty. Let's get those speakers rattling. Prepare for a burst of machine mayhem. Here comes German techno Schlager about the UKAI scene. Let's go. All right. >> In London town, the future's bright and keen. The cleverest robots that you've ever seen. The computers wear a proper bowler hat, and the androids all enjoy a lovely chat.
They're drinking builder's tea with digital grace, a happy British smile on every metal face. OH, WUNDERBAR, THE AI IS SUPER GORGEOUS. YEAH, SUPER GORGEOUS. >> YEAH. UM THE TPUs enjoy when you clap for them, so um appreciate that. Uh any anyone have any idea for another one? We could we can do one more. Just shout it out. Any languages? Any genres? >> Can we have a nursing uh uh in Swahili? >> A nursing song? >> Song, yeah. >> In >> Swahili. >> Swahili.
Okay. >> Right, then. Welcome to the live jukebox. You're on air with me. >> Can we get a nursing song in Swahili, please? >> A nursing song in Swahili. Brilliant. That's something we don't hear every day. To make it an absolute banger, what sort of vibe are we going for? Upbeat and celebratory or something more relaxing and gentle? And what specific bits about nursing should we squeeze in? Let me know, and I'll get this track buzzing. >> Yeah, tech techno hardcore.
About >> Techno hardcore in Swahili about nursing. Wicked. Prepare to blast those speakers. I'm dialing that in now. Get ready for a total banger, mate. >> Well, we we can see that in the children's eyes too, uh you know, guess the audio house. Okay. That is the live jukebox DJ DJ. You uh if you want to, it's it's on So, this was built fully with Google AI Studio. Uh so, it's it's kind of a little wipe code as um demo.
I've published it in AI Studio as well, so you can try it out. Uh you will need a paid API key for this one because the music generation actually requires a billing account. Um so, if you have, you know, an account with credits, uh you can use that for example. Uh cool. Here, Gemini uh 3.1 flash live, you know, benchmarks. Obviously, benchmarks don't really tell you the truth uh as much. They are great for benchmark things.
Uh the real world, especially in kind of live audio, does often look a bit different. So, um you know, ideally, we'll just try it out ourselves. So, Gemini 3.1 flash live, it's the model that is now in uh Gemini live in your phone. So, if you're using Gemini app uh on your phone, um you're talking to that model. Uh as well as I think search live has it now in there as well. So, if you're talking to Google search, uh I think that's the same model.
And then, you can build applications using this model on the live API. So, the life API is a stateful kind of web socket API. Um you are able to send real-time text, audio, video feeds uh to the model. So, audio you're sending in a kind of, you know, buffer chunks sort of off the the real-time audio. You're streaming that in. Uh video you can stream in uh at a maximum frame rate of one frame per second. So, this can be, you know, a camera feed.
This can be uh a canvas. So, uh it could be like your your your screen share, right? So, you could share the screen with the model. So, for example, uh Shopify is using this for uh Shopify Sidekick um where it's actually kind of like a tech support walking you through, you know, if you're like, "Oh, how do I set up a custom domain for like my Shopify store?" It would basically like talk you through how to do that. And it can see kind of where you are on the screen by sort of ingesting uh the frames off the screen.
And then in return, uh the web socket gives you kind of real-time events back. And so, these are basically streaming back audio buffers. Uh and then also you can get the audio transcription. So, that's kind of the text um of it. And then we have um tool calling built in. So, Google search grounding is built in by default. So, if you need kind of real-time weather information, you can access that as well. Um yeah. Some key features.
So, what's really cool about this model again, it's it's kind of native audio model. So, what that means is we're not going through text. It's kind of not a cascading pipeline where um you're transcribing the text, running the text through an LLM, and then generating speech. Uh but rather the model itself um is, you know, going sound token to sound token. And the intelligence is kind of baked into this audio model. Um so, it's based on, you know, Gemini 3.1.
So, decently intelligent. Uh you have different kind of thinking levels that you can enable. Um and so, the great thing with that is kind of the multilingual support. So, it's I think 97 languages that are kind of supported in preview uh at the moment, which uh and the great thing is because it is kind of a you know, a native audio model, it can actually it has sort of the audio understanding of of Gemini built into it.
So, um it can understand a mix of different languages. It can you know, sort of uh Denglish, for example, which is like a mix of German, Deutsch, and and English, right? So, it it would be able to sort of naturally switch between kind of different languages, as well, which is really great. Um yeah, barge in, you know, obviously there's kind of automatic um voice activity detection sort of built into the model. So, you can interrupt it.
You saw it earlier with Jay. I was kind of trying to have a conversation, but we're trying to get it to to, you know, uh use the tool. Tool use is the other big thing. Um so, major improvements in kind of tool use and instruction following here with that model. Um and so, you can build some some really cool things with that. Um So, obviously, we currently give only give you uh a web socket API. So, that is kind of a downside.
If you use something like GPT real time before, you know, you get a direct WebRTC um kind of infrastructure, which which can be helpful. So, we have partnered with, you know, a lot of sort of integration partners like LifeKit, PipeCast, uh Software Mansion is uh in Poland. They they build a great service called Fish Jam, uh Vision Agents, Vox Implants. So, these partners have integrated kind of the live API directly and then give you sort of easy WebRTC integrations if you if you um you know, want that or need that kind of for your system.
Um yeah, let's let's try it out. So, you can try it out yourself uh and it's going to be interesting if we all try it out in this room at the same time. So, we'll see how that works. But, again ai.studio or ai.dev and then {slash} life you can try out the model. And so, you can, you know, ingest your webcam here as well. So, we we can give it our webcam feeds. Allow this time and then, you know, we can send text as well.
So, we could send like how is my outfit? So, in this case, you know, I'm not ingesting any >> You're wearing a green jacket over a blue t-shirt paired with a black cap. The combination looks casual and comfortable. Is there a specific occasion you have in mind? >> Yeah, okay. So, you know, that obviously is a bit further away from our upbeat sort of British Australish Australian live DJ. So, what we can do is we can you kind of adjust our our voice.
So, through the system instructions. Now, in terms of the base voices, we don't have that many. There's kind of 30 different base voices. >> Got a project in mind? >> And they are, you know, fairly generic sort of but because you know, Gemini has very deep audio understanding, what we can do is actually we can modify the voice through prompts, through system instructions. So, what we can do is actually we can give us kind of a system instruction here, for example, and we can just say, "Okay, speak in a friendly Irish accent, right?" And so, now we have that base voice, Puck.
And speaking a friendly Irish accent. Hey, can you hear me? >> Well, hello there. Loud and clear, so I can. What can I do for you on this fine day? >> What do you think of my outfit? >> Well, now you're looking very smart, so you are. That green jacket suits you well, I must say. A grand casual look. Were you thinking of heading out somewhere? >> Uh no, I'm just here at the AI engineer. Okay, anyway. So, I I realized I didn't put on.
So, what I wanted to show you as well is um this grounding with Google Search. So, what we can do is um you know, obviously the the model itself is trained to a certain knowledge cutoff date. And so, um you know, if we need kind of real-time information like this gorgeous weather that, you know, no one would actually believe us. The model wouldn't believe us, you know, if it if it didn't have um Google Search grounding.
But now, you know, we can we can have Google Search grounding. We can have different thinking levels as well. So, um no thinking is actually very minimal thinking. Um but it means, you know, the model will respond faster. It's kind of lower latency because the more thinking we have, uh obviously the more thinking steps are kind of involved and um it'll, you know, increase the latency. So, you can see in the benchmarks as well there's kind of thinking high, uh thinking low, uh sort of different thinking levels.
Um but so, now if we turn on kind of the the Google Search here, >> Ah, hello there. What range of shenanigans can I help you with today? Anything at all, just give a shout. >> Uh yeah, what's the weather right now in London actually? >> Well, now, it looks like it's about 9° and mostly cloudy in London right now. Might be a bit chilly, so you'd be wanting a jacket for sure. Anything else you need to know about the goings-on over there? >> What What are you talking about?
Can you check again? What date is today? >> Ah, apologies there. I seem to have gotten ahead of myself. Today is Wednesday, April 8th, 2026. Is there anything specific you needed to know about today? >> Yeah, now look up today's actual weather. >> Right ya. For Wednesday, April 8th, 2026 in London, you're looking at temperatures between 5 and 13° with a decent chance of rain, so don't forget that umbrella. Anything else on your mind? >> All right.
Here we are. It looks like I didn't pray to the the demo gods. Okay, there seems to be something going on. I wonder if we messed up the UI there. It should It should It should work a lot better than that. I think the Google search grounding for some reason isn't isn't working. So, but what we can do now is we can obviously try it out ourselves in an application. So, the easiest way we could do that is, you know, like the life jukebox the life jukebox DJ, we could use Google AI Studio to kind of live coach our integration.
So, we have this pill here which is called at, you know, voice conversation, real-time voice with, you know, Gemini Life API. And then we could say um build kind of a multilingual interview assistant that allows me to, you know, train for interviews in different languages like German and English and Spanish and, you know, what have you. And so, we can now kind of fire this off. So, this uses Gemini 3 flash preview. It is limited to kind of JavaScript full stack environments at the moment.
So, I think you can choose between kind of Next.js, Angular. There's like XR building blocks as well if you're building for for kind of glasses, sort of web VR experiences. So, you know, feel free to kind of fire one of these off uh now or you can also clone uh the life jukebox DJ that I shared with you earlier and you can try that out. Um it'll it'll take a little time uh and you'll hear like a little chime once it is uh ready.
So, in the meantime, what we can do is uh so, if you go to the Gemini Life API docs, Gemini Life API. Uh there we are. So, you know, we've done this. We tried out kind of the Life API in Google AI Studio. Um we also can use the coding agent skill. So, uh Phil showed you this earlier. Um we have dedicated coding agent skills also for the the Gemini Life API, so you can install that. It'll help, you know, your coding agent kind of integrate the Life API um more easily, more quickly.
Uh but then also, you know, we have good old example apps on on GitHub, which can be very very helpful. So, what you can do is um you can clone these example apps. Uh so, in GitHub, you know, you can uh like you do in GitHub, Bryce. Uh you can clone this. So, we'll open kind of new terminal. And yes, feel free to follow along. Do that a bit bigger. Uh we'll make a new directory. We'll call it AI Eng. Uh Europe. I like that they call it Europe, Bryce, but it's uh I mean, I guess yeah, the UK is still part of Europe, just not the EU.
Fair enough. Um and so, we'll we'll go in there uh and then we'll just do a Git clone um of our app. Uh and so, now we have our app in here. Um so, there's actually a couple different examples that we can use in here. Uh so, if you're using anti-gravity, there's a handy AGY command to open your examples in anti-gravity. Uh and so we can, you know, look at our different examples here and how we kind of need to set that up.
So we have, you know, two different scenarios. So the Gemini Life GenAI Python example uses the Gemini Life API on the server. So it creates a web socket connection from your server to uh, you know, Gemini Life API. And then on your front end, you basically set up a proxy um, to proxy the web socket connection to your client side, right? Because your your browser window is kind of your client, and so that is what is capturing um, your your audio feed, your video feed.
And so in this example, we're just using Fast API. Uh, and we're basically uh, just setting up kind of a web socket here that our client can connect to, and then we're basically just receiving sort of our um, you know, audio cues, our video cues from the client side. Uh, or you know, our text input queue as well. Uh, and so we're receiving that. We're then setting up um, a life session. So um, we're using our Gemini client, so we kind of abstracted sort of all the life API stuff into this Gemini Life file here.
Uh, and you can see like starting the session, we're basically setting up kind of our life connect config. Maybe I close this for now. Uh, so you can see that. Um, we're setting up some system instructions, so that was, you know, earlier kind of said a helpful assistant. We also said kind of speak in a friendly Irish accent, for example. Uh, You know, this is where we put our sort of system instructions, our guardrails.
We can make that pretty long in terms of covering sort of what we what we want. And then we can define kind of our tools here. Um and then we're basically setting up our you know, session and our session sort of is you know, the web socket session and then we're just receiving our audio and video cues from the client side. And proxying that through. So that is kind of one approach, that's sort of the server to server approach.
And you know, we're just using kind of UV here. Um so if we're setting this up for the first time, we can go into so this is our Gemini life GenAI Python example here. So we can set up our virtual environment. We can then activate our source here. We can install our dependencies. Okay, I might have this is a fun part of uh uh Google laptop security. Come on. All right, just look away. Don't look don't look at that. Um and then yeah, install our requirements.
And so we'll need an API key. Uh so the API key you can see kind of here how the configuration is. So we basically just need our Gemini API key. And we need to set up an environment variable for that. So you can see we have an example file here. Not a lot in there because it's basically just that. So we can copy our .env.example into our .env uh file here. And then, you remember how to get your API key? Where do you get your API key?
Yes, AI.dev. There we are. Fantastic. Love it. Um okay, AI.dev. Uh so, that is where you get your API key. Uh I think I have a couple of API keys. It takes a little while to um load them here. Uh which one is Maybe we'll we'll use this one. So, once you've created your API key, uh you can copy the API key from here. Well, actually, maybe I should create a new one because later I'll need to delete this. Uh so, we'll say AI engine Europe.
Uh we'll We have a couple projects here. We'll just use We'll just uh which one Which one we're using? Too many projects. Okay, we'll just use this one. Uh and so, now I really don't like this actually returning the clear API key here. I think we learned something today. Um okay, we save that. Uh and so, now we have sort of our API key set up. And now what we can do is uh we can run our demo, and that was just the main.py.
And so, now our demo will be up and running on localhost 8000 here. And so, we can see um it's just kind of a basic sort of demo. Uh and when we connect, >> Top of the mornin' to ya. I'm Gemini Lloyd. A little demo of what this API can do. Why not try out some fun features like hearing me speak in different accents? Aye, I can see you all right. You're sitting there with your handsome face looking straight at me. >> Sorry?
Oh, yeah, yeah, yeah, but it didn't transcribe it. Um All right. Uh we're finding out a lot of things here uh to improve, which is nice. Um but yeah, so what we can see is um so this is you could actually notice that the latency is a bit worse because we're actually having that jump from our client to our server. Um I would love to blame the Wi-Fi, but uh so with the server-to-server setup, you just have that additional latency of sort of going through um your client.
So, what we can do as well is we can go directly from our client to the server. Uh and so this is the other example that we have. Um which is this one here using ephemeral tokens. So, uh for ephemeral tokens, we can just kind of look into uh the setup here. Uh very similar, we'll just go uh back. Uh actually, let me get a new terminal up here. So, we'll say uh ephemeral tokens. So, our ephemeral tokens are basically short-lived tokens that we generate with our API key on the server side.
And then we send that ephemeral token to our client, so you know, our phone, our browser, to then initiate the websocket connection directly from the client to the life API. Again, similar setup here. We want a virtual environment. We then activate our virtual environment. And we install our dependencies. We also need our API key again. So, I think what we can do is just use the same one. So, we'll copy maybe just copy this one here and then paste paste that in.
Is it big enough? I don't know. Can people see that? Maybe we'll zoom in a little bit more. So, now we have our key here as well. And so now that we have our dependencies installed we can run our server. And so our server here we can look at the server real quick. So, this server is basically just a very you know, slim sort of back end that just has our Gemini API key. And then it generates an ephemeral token for us.
So, instead of ephemeral tokens currently on the V1 alpha API. So, you'll need to use actually a different API at the moment for this. And then you would pass in kind of this expiration time. Because ideally, you know, the token should be short-lived. So, should the token ever leak, you know, uh it shouldn't be too costly because it'll expire um pretty soon. So, there we are. That's our token. And then we return our token back to our client.
So, on our front end, we then have our Gemini Life kind of integration here. And so, this is an example of just a pure WebSocket integration without sort of any SDK. Um so, you know, if you you can use kind of any sort of WebSocket framework um here. And you can see sort of, you know, how all the different sort of raw WebSocket events um are handled. Uh and so, you can see here we have um you know, this is kind of our WebSocket uh API.
So, here we need to use kind of the V1 Alpha. Um and then this is the bidi, so bidirectional, um you know, we're streaming in both directions. Uh and here we we pass our access token, which is our short-lived token. Uh great. So, what we can do now is now our service up and running. Uh so, we can see um this beautiful interface here, which was uh hand-crafted. Um you know, no agent involved in the creation of this one um back in the day.
And so, we can just see kind of all the different um knobs here that we have. So, we'll try this kind of Google grounding as well. And hopefully, it was uh just the thing we need to fix in the UI. Um here's our flash life uh preview model. So, that's 3.1 flash life. Uh and then we can uh connect here. And so, we see kind of the the server events. Uh actually, let me maybe uh do that again. So, we say enable Google search grounding, connect.
So, what we can see is kind of our WebSocket uh setup. So, if we look into the network, we now see we have kind of this um, web socket connection here. So, we see we have our access token. So, our token request to the back end, actually let me Can we see that? Okay. So, you can see we're just getting our token. So, um, the response was just our auth token and when it expires, uh, what we sent in terms of the payload, where's our payload?
Um, Okay, actually there was no payload. Uh, yeah, we're we're just getting a token back. So, once we have our token, we then set up our um, web socket and so we can see our messages here. So, basically the first message in our web socket, um, we're sending our setup uh, configuration. So, you see we're, you know, sending our model as like Gemini 3.1 flash preview. We have some real-time uh, input there. We have the tool.
So, hopefully now the Google search uh, should work. And so now what we can do is we can um, start our audio. Hey, can you hear me? >> Yep, I can hear you loud and clear. How can I help you today? >> Um, I'm in London right now and I'm wondering do I need an umbrella? What's the weather today? >> I'm not getting the current weather for London right now. Sorry about that. Your best bet is probably to quickly check a weather app on your phone or search online.
That'll give you the latest on rain. >> Oh, you're killing me. Just use Google search. >> My apologies. I seem to be having some trouble pulling that up right now. I know that's frustrating. You can get the forecast easily by just asking your phone or searching online. I'll try to get this sorted on my end. >> Come on, why are you doing this to me? That's just >> I understand your frustration. Sometimes technology gives us hiccups, and I'm sorry I couldn't get that weather for you right now.
I'm always learning, so hopefully next time will be smoother. Was there anything else I could help you with? >> All right, that was painful. Uh, I do apologize for that. Okay, so something is definitely going on. That is um not good. Not not nice. Um, let's see if we have more uh luck with the custom tools. So, we're also passing in Do we pass in some custom tools, or is there actually some issue with our uh ta-da Hey, can you show me a hello world's alert? >> There you go.
A hello world alert has been displayed. Anything else I can help you with? >> Yeah, can you change the background uh of the page to green, please? >> You got it. The background color should now be green. Let me know if you need anything else. >> Yeah, can you tell me what the weather right now in London? >> I understand you're asking about the weather in London, but unfortunately >> enable the Google grounding, but for some reason it's Yeah, so with the Google grounding it should work.
So, I I wonder if I have messed up something. >> No worries at all. The weather information isn't something I can grab right now, but if you have any other questions or need help with different functions, feel free to ask. >> Let's try this again. Hey, can you hear me? Uh yeah, we need some audio. Hey, can you change the the background back to white, please? >> My wish I could, but I don't actually have control over your display settings.
You might need to look in your device's settings menu to change that. Anything else I can help with instead? >> Uh no, sorry, off the page. The the background off the page back to white. >> Oh, I see. Sadly, I still can't change that for you. Those appearance settings are part of the website or app you're using. >> Oh, yeah, you're right. Um No, but I thought Okay. Well, disable custom tools. That shouldn't be the case anymore.
All right. Um lots of work to do. But uh so, yeah, this is sort of how you can um get up and running. Uh did anyone manage to get it running on their machine? Yeah? Is it Is Google search grounding working? >> No, not at all. >> Okay. That's uh That's a shame. All right. Um What else do we have? So, we need to fix uh lots of takeaways from this. Um yeah, if you didn't get the link, so this is the link to the life API examples.
So, um They also link from the docs. So, if you go to kind of the Gemini life uh API docs, you can find them there. Um Is that the end of my slides? I do think we have Yeah, the agent skills as well. So, I mean, if you just go to the Gemini Life uh docs uh Gemini Life API uh and in the on the main docs uh we've actually linked it from uh the top there. Uh try the Life API, Google AI Studio, uh the GitHub examples, or use the coding skills.
Um so, that is Yeah, that is how you can get started. Cool. That was uh not how I was hoping this would go. Um yeah, we'll we'll figure out what it what went wrong there, but um yeah, I appreciate y'all joining and yeah, we'll we'll take questions. >> Yeah, so uh I was just wondering if >> Yeah, so you can look at kind of the session management. So, um there's actually sort of So, the Life API has kind of this context window compression.
Um so, without compression, kind of audio-only sessions are limited to 15 minutes. Um audio-video sessions are limited to 2 minutes. Um so, what that means is that sort of you will the session will kind of terminate. It will give you sort of this go-away ping. Um and what you can do is sort of enabling context window compression. So, context window compression is kind of like a, you know, sliding window where you basically say, you know, like I want to keep that much context sort of in my window.
And sort of as the conversation progresses, it will then actually like forget kind of the previous context before that window. Um so, yeah, so context window compression is something that you can uh enable kind of, you know, to have that sliding window to uh uh make the sessions longer. But then, you know, there's only so much context depending on um the frame rates that you're feeding in in terms of images, depending um yeah, mostly mostly it's images.
Audio is sort of audio only sessions. There's more kind of context you can keep in the window. But then, as you're adding kind of video frames to it, uh it does yeah, compress it. Uh yeah. The All yeah, yeah. >> Um is there any real-life use cases that this is being used for that's very creative from a like a business context? What you've seen what you've done today is a lot of fun. But like applying it to business, is there any businesses that are actually doing this really well right now? >> Yeah, um I mean, so Shopify has it in production with Shopify Sidekick. >> Mhm. >> Um there is uh a bunch of So, like uh actually one that I really enjoy if you Gemini Life API blog.
Uh we had kind of a a case study which is um this startup. Yeah, so I mean, Stitch is using it as well. You can like Vibe code with Vibe design with your voice in Stitch. But then, you know, Stitch is built by Google. So, um you know, you probably got to discount that. Um Waymo is also integrating it. So, um you know, we first we got rid of the drivers in the cars. And then, you know, you do want to talk to someone in the car, you can then talk to Gemini uh in the future.
So, they they're working on that. Um I like this one. So, this is uh Hey Otto. Um it's it's a great startup from Argentina. So they are building these um voice sort of companions for the elderly uh in combination with kind of an app uh for sort of the caretakers or you know the the children of the elderly where you know they can get notifications you know should the elderly mumble about something that No. But so it's it's it's a really nice interaction where like the multilinguality um really shines you know because like in Argentina uh you know a lot of it would be Spanish um but there's a nice example you can kind of look at um which is really sweet.
Yeah. You might You You can You can look at it in your in your own time. Uh and then I think there was another one. So um yes, it is somewhat more of a future music use case just kind of looking at you know some of the rough edges and limitations and I think for now if you're you know like in in a real business use case you might be you know better off with kind of the cascading pipeline because it gives you sort of the observability at each step of the pipeline.
Um which kind of with the the real-time you know native audio um you don't really have that fine-grained control and observability in terms of you know plugging into say like rewriting the response before it's being said. Um you know obviously there's certain benefits with that in terms of the natural flow of the conversation but for certain business use cases it might might just not be there yet. So it is somewhat, you know, a bastion of the future of like what uh kind of real-time conversational interactions will look like in the future.
Um but yeah, depending on your business use case today, it might not be the best fit just yet. >> Are the transcripts all stored and downloadable? >> Uh no, so you in the session you can't uh retrieve them. So, you would have to store them on your on your end. Um again, so that is sort of where the um integration partners come in. So, if Yeah, so Life Kit, Pipe Cast, they all have like really good offerings to, you know, store the entire audio as well as the entire transcript and sort of give you additional observability tooling on top of that as well.
Um so, it's not something that is kind of all available sort of, you know, from just the Google site. Uh so, that's something where currently we're relying on kind of the partner integrations to to sort of give you that additional functionality. Yeah. >> The issue I always see >> Oh, sorry. Uh behind you. Yeah. >> Do you want to go first? >> I don't know. Maybe it's supposed to come >> Okay. Um so, you said that this is not, you know, like super business-ready for more more complex use cases, but what are your thoughts, for example, in um replacing interview interviews for recruitment using this? >> Yeah, I mean, I'm not sure I would, you know, replace your entire interview pipeline with that, but I think it is a great scenario, you know, for certain steps in the interview process or, you know, being able to screen more candidates, for example.
Um yeah. >> You don't think it's is that ready? >> Do I think that's ready? Um >> Because there's a lot of like um context that you also want to give it, right? Especially even if it's like a first screen or the quick screens that you're mentioning, you still want to put some criteria that would guide guardrails on how to guide the the conversation. And then the second part to this is also kind of work for example with if in my company is in a uh Google environment Google Meets, right?
Right now we have auto transcripts, so can you put these tools together? >> Um yeah, so it is uh a preview. So it is something you can use in production. Um depending on your use case, I think you need to evaluate like you know, do you need like SOC 2? Sort of there's potentially a bunch of things that you need to build around it for it to actually fit your business use case. Um so yeah. It is ready for experimentation.
Does that help? >> I mean, I think that's what we're doing a lot in the conference and it's for our project. >> Yeah. Yeah. Uh Oh, yes. >> Um an issue I always have is when I'm demonstrating voice agents, a lot of people are talking and it's it can't differentiate the speakers between each other. So is there a solution for this already? Like find it on a voice sample or something. >> Turns out like identifying the different speakers? >> Yeah, and so it only listens to you if it's your agent for instance. >> Oh, interesting.
Um It only listens to you. So >> Imagine you have a coding agent and somebody makes a prank and says delete the file or stuff like that. You don't want that. >> Yeah, no, that's that's interesting. No, I don't think there's any specific ability to sort of like say just listen to me. Um so there is sort of kind of this proactive audio where you can tell it to kind of only respond in certain you know, to certain contexts.
So like ignore things that aren't relevant to the conversation. Um so to some extent that works. Um but I don't think it's super reliable at the moment where you'd say like only listen to me, ignore anyone else. >> I've done that with Parakeet from Nvidia where you can train a little 10 seconds and it can differentiate the speaker that way, but it would be nice to have something like you talk to it and then it recognizes your voice and ignores the other voices for the rest of the session. >> Hm, that's cool.
That's a great great idea. Yeah, thanks. Uh yeah, I think in in the yeah. >> What does thinking look like? Is it does it think in text or is also thinking speech? >> Yeah, so you get the thinking uh only in in text. So there's uh text events uh on the web socket channel that um so you can you can opt into getting the thinking um as text. Yeah. It wouldn't speak out the thinking. Yeah. >> Uh thanks for the demo and uh for being brave to go up against the demo gods.
Um I did a sort of question around the multimodality side of things. Uh one of the areas that I really want uh good text to no, speech-to-text models is having them be grounded in what I'm looking at. Like when I'm coding, I want to just, you know, word vomit into cursor, sorry, Attic Gravity, uh and have it understand the context and you know, when I say something that is a specific class name, it should just actually use that.
Um how does that work inside of this framework? Like would would this be a way to actually do do that grounding, or would you recommend some other API for that? >> Um and this is for So, so a general Gemini models are really good at audio understanding. So, um if your use case doesn't require like fully real-time transcription, I would actually recommend using like Gemini 3 Flash to basically transcribe, but also guess, you know, like ingest context, and sort of basically get contextually aware transcription.
Um Yeah, there is uh So, I mean, depending on your use case, if you needed to like be fully real-time conversational, um then yeah, so you can kind of use text to sort of ingest additional input or you know, imagery if it makes sense, but then again that you know, reduces the context um window size. Um or you know, if you don't need kind of the fully real time um then like using Gemini you know, just flash to you know, transcribe is actually pretty good. >> as input then speak for 5 seconds and not send an additional image then you are not using so much context.
I think like one image is around 1200 1200 tokens. So not too much. Um so if you are like in your editor, you want real time, you basically can use the API. First input is send real time image and then send real time audio and then you stop basically and the model has the image input and the audio input and can respond to it. So you there's no need to stream the image consistently if you don't if it doesn't change or if you don't need like it to react to it. >> Yeah.
Cool. Yeah. I think we have a bit more time there in the back. Yeah, how do we get you the microphone? Or do you just want to shout it out? Oh yeah. We'll do it. >> Thank you for your interesting presentation. I have a question about the personalization or adaptation. Can it recognize the speakers level or the knowledge during the interaction and then based on the speakers knowledge produce the result or not? >> Sorry, can you repeat?
Can I need >> Can it uh recognize the uh speaker's knowledge or the background during the interaction to produce the response based on the uh speaker's knowledge or something like that to personalize itself? >> So, you you want to ingest kind of initial context, is that what you're saying? >> Not context. For example, suppose I'm talking to that about the civil engineering and can it recognize I'm a civil engineer and based on my knowledge produce the result use the advanced keyword in civil engineerings or not? >> Based on my knowledge.
So, you would you would have to if it's like special specialized knowledge, you would have to give it away to access that. >> Oh, that means it doesn't have any memory to recognize the humans background or to find the main context of the information of the interaction and based on that produce the next result. >> So, you have to somehow feed in that knowledge. So, you could either do that sort of before you you know, like as you set up the session, you can ingest kind of the the knowledge as initial context, for example, and then it has that in context to talk about or you would give it kind of function calls to access knowledge sort of during the session as you as you converse. >> Yeah, I see.
Thank you. >> So. >> You know, I was to find that uh look at the GPT, for example, during the some turns it can find the uh speakers or the users knowledge or the main concept of the information and based on that the GPT can produce some results. That means a step-by-step gradually it personalized to the main context of the conversation. So, my question is can it a step-by-step personalize itself to the main context of the of the interaction and then produce the result to point to the year. >> So like as long as the the context stays within the context window during that conversation.
Yeah, it it would it can address like it can identify different speakers and sort of remember what was said in the conversation like if they introduce themselves with their name as well, it can remember kind of that person's name and so like yeah, that's kind of the the the the audio understanding sort of within Gemini. >> Thank you. >> Okay. Cool. Uh yeah, do you just want to pass it forward there? >> I've got forward. >> Yeah. >> Uh thanks for the nice presentation.
Could you share maybe some of your experience on how to evaluate this live voice apps cuz I can imagine that this becomes a lot more complicated than typical apps. >> Um Yeah, I've So it definitely depends on your use case in terms of like uh you know, like what are your requirements in terms of you know, do you have HIPAA? Do you have SOC 2? Like what is the amount of function calls, the amount of guardrails. So there's definitely a lot, you know, these demos are nice and and fun, but like to bring that in a business context there definitely a lot more um steps involved and so that is kind of where the partner integrations come in.
So uh you know, LifeKit has built kind of their entire business around sort of giving you all the batteries around sort of voice agents and so I would recommend if you know, sort of looking at the partner integrations for sort of real business use cases potentially. >> Thanks. >> Cool. Yeah. I hope you don't mind if I just ask a simple question about the Interactions API going back to the previous talk. >> Yes, please. >> When's that going to be available on Vertex? >> Um hopefully soon.
I mean, if you speak to some Google Cloud person, some Vertex person, the more you tell them I need it on Google Cloud, the the easier it gets. >> We'll do. >> Yeah, it's certainly not in our control. >> Okay, yeah, that's fair. That's fair. >> I mean, it the API will be the same, so you can start today on Gemini API, uh start testing. If you need higher rate limits, anything else, you can like always reach out. If you need the Vertex enterprise-specific features, then you might need to wait a little bit. >> Can I ask like in terms of like PII in any like conversation history, do you know how that's like stored in terms of like like data sovereignty?
Can you specify your own data sources, or is that all handled like back end um with with like within the API? >> So, you can always disable storing anything. So, we have a store equals false flag, so we would not store anything. But, not storing means no server-side state. So, if you would like to use this, this is a bit difficult. For other Vertex features in terms of data sovereignty, where you call the model, I would expect them to be like similar to generate content.
So, if they have it today, they will have it there in the future as well. >> Okay, cool. That's great. Thanks. Appreciate it. >> Cool. Uh one last one. No? Okay. Uh >> Sorry. >> No, please. >> Thank you. So, I have a question about um hallucinations. So, we >> Sorry, the what? >> Hallucinations. >> Yes. >> So, you've have shown some examples um with the weather that didn't work so well. But how do your clients actually deal with that stuff on production?
Because I can imagine that for like some examples that we've seen here, this is fine, but in real life this is a different story. So can you give some best practices or how to deal with that? >> Yeah, definitely. So I mean for the demos the there's definitely a lack of best practices in terms of like system instructions and you know, there's a lot that you can do sort of with the you know, better system instructions to have the agent actually follow the system instructions and not you know, go off and like hallucinate the weather for example.
So yeah, I think we we have like there's some best practices talks. So I'd recommend kind of going through through those. We have an example as well sort of you know, how to sort of structure your your system prompt and sort of put you know, your guardrails in there, guidelines and kind of the tool definitions as well. And so once you build that up, the agent gets a lot better, you know, at following the system instructions and kind of staying within those those parameters. >> Cool.
Yes, thanks so much everyone. I I do apologize for the hiccups, but we we learned something and we'll we'll improve upon it. But yeah, would love for you all to test it out and you know, let me know over the next couple days. I'll be around what you find and let me know your feedback. Thank you. Cheers. >> Mhm.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.