Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
Most replayed moment at 13:49
2.3x that video's typical replay level
Fable uh and it runs into an unknown, ask it to log it, right? So that um you uh you can see where the deviations happened and then you can sort of figure out why as well, you know? It will usually give you some context about what happened.
Said at 13:43
The graph counts replays. It does not show where viewers stopped watching.
Words
15,181
Runtime
1:38:34
Speaking pace
154wpm
Reading time
63min
154 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Good afternoon everybody. Welcome to sunny London, hey. Um is everyone's first time here? I think it is because this is the first conference. Amazing. Amazing. Well, thank you very much for joining today's session. Um hopefully you are in the right session but for those who uh need to double click. Uh this is a hands-on workshop to delivering quality AI applications with BrainTrust. And we'll also be partnering with our colleagues at uh Trainline, which
77 words, the words spoken in the first 30 seconds at 154 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 1,022 |
| Average words per sentence | 14.9 |
| Longest sentence | 63 words |
| Questions asked | 52 |
| Sentences containing a number | 18 |
Most used terms
Filler phrases
895 in total: uh 400 · um 243 · you know 80 · like 66 · kind of 58 · actually 24 · right? 10 · basically 6 · sort of 6 · I mean 2.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
Good afternoon everybody. Welcome to sunny London, hey. Um is everyone's first time here? I think it is because this is the first conference. Amazing. Amazing. Well, thank you very much for joining today's session. Um hopefully you are in the right session but for those who uh need to double click. Uh this is a hands-on workshop to delivering quality AI applications with BrainTrust. And we'll also be partnering with our colleagues at uh Trainline, which I'll introduce myself shortly.
Uh so, probably you'll be wondering who's this guy? Does he even go here? What's his name? Introduction to myself is uh you can call me Duran, bit like the band Duran Duran with a G, so hopefully folks are saying I'm a Duran Duran fan. And uh I've had a bit of part of helping organizations and enterprises how scale adoption of mission-critical systems and now moving into the age of AI. Um you know, background in uh formal mathematics, so I know there's all the rage going from machine learning to data science and now AI engineering and this comes at a very topical time.
Uh feel free to connect with LinkedIn if you want to stalk or just want to uh you know, just general chit-chat around this area. I'm also joined by my two friends over from Trainline. If you want to come over and introduce yourselves, team. >> Of course, if the mic is working. Are good? Yes. Hello everyone. Uh thanks for coming to this workshop, especially after that lunch break and uh uh and the sunny time outside. Thank you very much for coming here.
Uh so, my name is uh Osama. I'm a senior AI/ML engineer at Trainline. Uh for those who know me, yeah, one person in the room. Uh yeah, I was staff platform engineer before and I have uh background in computer vision um and I was doing also mobile apps on the side, so nothing to do uh with with AI at at some point, but yeah, now we are doing AI together with Mayank at Trainline. If you go. >> Sure. >> Hi everyone. My name is Mayank.
I um I'm a senior AI engineer at Trainline. I have background research background in LLMs, but my research is in the pre-big LLMs, so like if you've heard of BERT or the older LLMs. Um so we're continuing building a state-of-the-art uh agentic products at Trainline. If you've come here from abroad, I'm sure you must have used Trainline. If not, please do cuz it's a um it's the state-of-the-art app for um buying tickets and other things.
So, uh you're welcome and we're very excited to host you at this workshop. And uh yeah, we'll guide you through the uh the hands-on uh experience. >> Fantastic. I'm also joined by some of my colleagues as well from Braintrust coming over the pond. So, Phil, Eric, and Rose, if you just show your hands up. Don't worry, folks, you're in safe hands if you're stuck. So, just holler and they can help. Fantastic. Uh just want to do a little bit of housekeeping as well.
If everybody could join into the AI engineer Slack channel of uh well, there's an AI engineer organization, but there's also a specific Slack channel which we'll be using today to help uh progress the workshop. So, if you are stuck, we can use the Slack channel to help each other out. Again, we're already on there, but as we progress and get to the hands-on element and if you are stuck, um this this will really help.
We're also providing a uh cheat sheet. So, if you are encountering any particular hurdle, there's a step-by-step instructions which help you just get to where we need to go to the workshop without, you know, having to to feel like you're you're you're falling behind. Again, a lot of the assets which we share today, uh you know, is publicly available. Um again, we can have any particular follow-ups um if needed. But, we'll give you a few minutes to make sure you get onto that and then join the channel AI engineer Europe 2026 Braintrust {dash} workshop.
It's a public-facing channel. All right, team. Um I we can start. Um yep, just for the people who are in the back. Um we'll need to you to join the Slack channel. So, if you see me here, we'll pick you So, that's it. Thank you, everybody. Um let's proceed. Okay. So, just kind of help orient today's uh workshop, we'll be breaking into kind of three main sections. We'll spend a little time to setting up a bit of establishing the background why we're here and and why this workshop is relevant.
Hopefully, we'll set the context for when we go into the workshop uh building uh the system, talking about this, you know, how do we ship uh AI equality. And then, we'll wrap up for uh the key takeaways, and we'll be around to answer any kind of uh open uh questions and answers that we have from the field. Okay. So, hopefully today um there's going to be a lot of people from very different backgrounds. Um we intended this workshop to be catered for Again, probably everyone here knows what an LLM is.
I don't want to insult anybody, but hopefully you're starting to uh explore your journey in terms of your um mature maturing your your operations in terms of building AI systems. So, whether you're an AI product engineer, probably an applied team or come from a traditional machine learning folks, maybe you might be a platform operation infrastructure, um this is really appropriate for you. Just kind of a show of hands here, like who here comes from a let's say a traditional data science background, perhaps?
Okay. Interesting, interesting. Who here perhaps comes from maybe software engineering or they're pivoting into AI? Okay. So, this is definitely the right room for you. Uh so, hopefully we'll we'll be able to kind of accelerate things as we progress. Okay. Um just to set the context here, um I think this is not an uncommon expression which we are seeing more broadly across the industry. Um again, put up a show of hands, who here has done a a machine learning or AI PO I see let's say let's say more specifically in generative AI POC, But then let's fail to kind of take that into production.
Okay. It's got a few hands here. I'll be worried for them if everyone didn't, but this is kind of a key thing which you know, speaking to me and my customers, you know, executives, top-level folks, all the rage, you know, they're thinking about this new technology. It's not necessarily new, but it's newer for a lot of folks, especially in kind of more enterprise and regulated industries. And then they're trying to take this to delivering value to their customers.
But unfortunately, there's a big hurdle between taking what you might develop locally on your machine and then industrializing it and making sure it works in angle. But the key thing that we see in from all of the research out there, it's not that the models aren't particularly smart. Like we've got very sophisticated models, whether you're building something in house or you're using kind of a top-off-the-shelf, you know, top commercial LLM providers out there.
What we do see more broadly speaking is the type of operational rigor when it comes to delivering these systems at scale has not kind of kept up. Because traditional software engineering, very deterministic, 1 + 1 = 2, great. LLM systems, you know, as my 10-year-old or sorry, 5-year-old would say, you know, 2 + 3 = 10, Daddy. So, it's having to kind of adjust it and make sure that we're delivering that to scale. And again, let's take a look at what we've seen when shipping these things is the fact that again, people think the demo state is suitable.
But clearly, doing two two to three five demos, great. Putting in production, everything goes awry. Treating some of your logs is observability and then this is really critical to to how we look at at BrainTrust is you know, logs will tell you what has happened, but sometimes you need to go deep into the system and understand its behavior. And this is really where observability comes into play. Something as well is you know, works on my machine, fails in production.
I try to patch the prompt. And then again, it's operational until the next issue happens or the next failure mode. But again, how do we keep track of that? If especially if you don't have a system in place, Um irrespective of what tooling you use, this can you know, has a categorical effect. So again, a lot of what we see again is not to do with tooling technology. It's down to more operational workflows. And this is really what we're aiming to help you in today's workshop.
Okay. So again, as I mentioned, it's not the prototype. It's getting to a state where we're knowing exactly what's changed in the system, how do we interrupt with that, and then how we systematically put a set of rigor so that we can get better and better. Remember, 100% our target is not 100% coverage. It's getting as close as possible while maintaining fixing the gaps that might have existed. Um again, something that we see time and time again is again, a one prompt might work, but as you move and industrialize it, you want to probably do things like breaking this down into each individual um sectors of responsibility.
So again, if you do come from the software engineering background, you know about these. You're breaking it down the moment into microservices. We'll outline a very similar approach here when we talk about building these systems. Um again, making sure that we're understanding these changes, uh putting a set of systems in place. This is really what we want to do in today's workshop. Okay. In terms of today, um hopefully, as I mentioned, this is a hands-on workshop.
So we're going to be going into the terminal, we're going to be going to the UI, and we're going to be going through the step-by-step and guiding you along the way. So we'll be doing a staged AI system uh with a multi-stage uh tool calling, which really allows us to um see this more uh let's say agentic flow. We'll then use BrainTrust to instrument and see how the performance of the application's working. We then also want to then, you know, take a look at creating what we call um or identifying a failure mode using uh a golden test.
So we'll push that through. We also then want to talk about how do we industrialize this. So then moving from, "Hey, it just works on my machine." to something that you can use in production and have it managed forward with a system in place. And key thing is identifying, you know, those those edge cases because again, you can create a a test data set, but ultimately, um there's no substitute for real-world data. So again, we'll be able to show you how we will will take those real signals in and evaluate and complete the loop.
Um so just a bit of introduction here today as well, who here's heard of Braintrust or played around Braintrust? Can I get a show of hands? Okay, great. Fantastic. So, okay. Love this. Love you all in there. So, just a bit of introduction to Braintrust, we're we're a company now that's uh I think just uh shy of 3 years old. Uh we're approximately a um a Series B company. So, we just announced that a few months ago. We raised $80 million at a $800 million valuation.
Um we have investors such as Iconic, Andreessen Horowitz, uh as well as Greylock, uh to really talk about, you know, helping organizations ship quality AI at scale. So, we're the platform for AI observability. Uh we've got a heavy user base uh globally, but we're also expanding our presence heavily in Europe. So, again, I'm one of the first engineers to to join and help build out uh our go-to-market function here. And I'm very excited again with our customers and our friends over at Trainline to do that.
Some of our other local customers include like Lovable, Doctolib as well. Uh we're still really pushing the forefront of these AI systems. Uh again, when it comes to using kind of Braintrust, again, I know to again, I want to kind of get your hands on with it, but where we really distinguish ourselves is being able to do it at scale. So, our founder, Ankur Goyal, um this is actually his third time building Braintrust.
Um so, he's an expert when it comes to a sort of database systems, and he's built um a a company called Impira previously that was acquired by Figma that talks about document extraction, and he's led the ML machine learning team out there. And because he realized, you know, building these these evaluations are hard. Understanding production traces are hard. So, if we're having this issue, I'm sure there are other organizations out there which are doing the same.
And so, he founded Brain Trust to really help do this at scale. And as we're doing and understanding this these these traces that are coming in and again being very highly semi-structured data which changes, uh he realized that traditional analytical systems weren't fit for purpose. So, we've put created a um a new category of databases called Brainstorm which really helps identify and accelerate this at scale. With us again, we're tool agnostic so irrespective of the agent framework you're using or the LLM providers out there, we we we're we're intending to help you deliver value uh irrespective of that.
Okay. One things I will talk about as we progress uh in this workshop is the there's a concept of a flywheel. So, uh again, if you ever come from Agile development, um you know, perfection is enemy of good. We want to start somewhere. So, even if it's the case that it's a new application and you don't know how it's going to be in production, we can start off with an evaluation set. If you have an existing application to to instrumenting, that's great.
We can pull that information in and identify the failure modes here. So, the key thing is get information into the system, identify those modes, remediate, ship it out, and then monitor, and complete the flywheel again and again till you get to where you need. All right, then. Um you would have uh a lot from me. So, one thing I'll do is I'm going to provide uh my colleagues over at Trainline to maybe just share their experiences of, you know, prior to to Brain Trust and and how they're helping us.
So, uh Osama, I'll give it to you. >> Thank you very much. Uh and hello again for people who are joining us uh just now. Um as my uh mate mine uh introduced Trainline, uh we are uh a company that actually helping people uh get on the trains. Um trains are different than planes uh if you don't know. Uh, there is a one like worldwide system, central system for all planes around the world. It's not the case for trains. Uh, and in Europe and the UK, it's very hard, I would say, if you would like to install an app for each career, it will take the whole space on on your phone mobile phone, for sure.
So, Trainline is actually being that platform, uh, basically, uh, to help you book tickets. Uh, mobile app agnostic, platform agnostic, career agnostic. You do it on one app and you can book a train from Paris to London, uh, from from, uh, from Lyon to to, uh, to Milano, whenever you like, basically. Every career in the U- Uh, we sell like almost like 6.3 billions, uh, of tickets on the trains. 27 millions active users and counting.
And the other interesting, probably, one, uh, for for this, uh, conference is how many, uh, AI conversations, actually, that we have from our with our travel assistant. So, we do have a travel assistant that is exposed to people and it's not just a chatbot. Uh, it's actually a multi-agent system that actually can handle refunds for you, uh, can handle changing trains for you. So, it is very, I would say, a proactive, um, agent system, agentic system and not just, uh, a chatbot.
Probably, Mine would like to add more about that. >> Yes, so one of the benefits of having 27 active customers is you've got a huge space of how you can serve agentic applications live to the customers. One of the examples of that is what Osama talked about is a travel assistant, which you can get to from a ticket window in the application. Um, it's an agentic system, which is something we want to talk about, uh, a little bit later in the slide. >> Awesome.
Which brings us, I would say, to the next, uh, point. Uh, I will keep it short. Uh, selling train tickets and being like, uh, a a a train uh, ticket company, how come that we are doing uh, machine learning? Of course, we we can. There's so much things to do and to help people in terms of their journeys, getting their tickets, getting their trains, getting back home, basically. We build We do two things. So, the classic ML part, which is actually building models.
We do that. We build ML models inside of Trainline from scratch, from data to model. This is something we do. And we also do the multi-agent uh the generative AI systems that we are now familiar with on top of LLMs that we love and cherish, all the toolings and context engineering and all of that. We do that. So, we do these both sides of the story at Trainline. And these are two examples of what users are actually using on top of those systems.
On the left side, so this is uh you can you can think of it as your weather application, but for train disruptions. So, basically, you have a ticket for a train, you get there. We know if this train would be would be disrupted or not, if it will be probably late or not. We know We know that, but based on huge data that we have on top of it that was a machine learning model that was trained and can actually predict train disruptions, being late, and all of that.
So, this is the classic ML part. The other one, that's the traveler system that I told you about. It is very As I said, it's very I would say advanced multi-agent system. It can show you alternative trains if your train is canceled or something wrong with it. Good luck doing that yourself, even with ChatGPT. And the the other one is handle refunds. So, you can actually get for a refund on your ticket if your train is late or all of that, and it actually can give you all of that without a handover, and also can you can do the handover to actual uh human customer support in our lovely customer support team.
Um which means that if you are doing this at production in production level and that scale at Trainline scale, it means that uh or you can ask the question, are we breaking things at Trainline? Of course, we don't want to to do that. Um and we are moving fast because technology is definitely is moving fast. And and this is why we are here. We are we're here to show you like what are we doing to move fast without breaking things in terms of course of AI.
And we can do that on top of handling the complex software systems that we we love and cherish cherish from APIs at scale and serving millions of users. Um so, how we would like to think about it is this scale. So, we know for sure that whatever we have our software systems, this is the deterministic side of the story that we have. And on the other side, building the ML models, this is the non-deterministic side of the story.
And we know for sure that the agentic systems are in between. There are parts of them that are deterministic. There are parts of them that are not definitely deterministic. And this is how we as the framework of thinking that we have. In terms of quality, how are we handling that? That's question. For the ML models, we do we do care about the quality of data and and the that we use for training the models. But also we have on top of it of course ML machine learning evaluations whether offline or online.
Offline it means like before going to production, you you need to do your evaluations. And online is the data from production you get it and evaluate your model if it's doing well or not. In our case for instance for that for weather forecast, uh is it predicting the state of the train if it's disrupted or not? Is it is the model correct in its prediction? So, that's one one example. On the other hand, for people who are familiar with software engineering, we do all, I would say, quality checks and the diagrams or whatever, and the tooling that we have for handling a quality at scale, very large scale of Trainline.
Uh and for those systems, I think you already guessed it, it's basically combination of both. It's not one without another, for sure. We do everything that we do quality-wise in terms of deterministic systems, but we also use the the email evaluation um side in from the non-deterministic systems. Uh so, yeah, that's I would say how the I would say the framework of thinking that we have at Trainline, thinking about these systems.
And um definitely uh BrainTrust is helping with that. And by the way, disclaimer, we are not I would say we are here just because we are convinced that it BrainTrust works for us. We we are not paid. So, that's 100% uh we we are happy customers. We we have been with BrainTrust for a long time. Uh and we use it in different um some of uh BrainTrust features, we actually use them here. So, for instance, for the ML AI evaluations, we do that and we we we follow the scoring of Travel Assistant on mini mini levels from uh from tone of voice to actual helpfulness um when it comes to ticket and tickets are really really, I would say, complex um in terms of reasoning.
And what you should get, what you should not get, depending on the if the train is late or not, uh the type of ticket is a return ticket or advanced, there are so many complex cases. Uh but yeah, we do follow that uh evaluation side. Probably My would like to add more about that. >> Yeah, so can I get a raise of hand from people who have struggled with LM costs, number of tokens, switching models, problems like this?
Right. So, yeah, it's nice to see. So, it was a problem with Trainline as well because we do this at scale and the amount of OpenAI and Anthropic amount that we pay is just like bizarre. So, we have to keep switching models, which is like the best model for our use case, cheaper models, talk to models with efficient tokens. Now, anytime you want to switch models, you want to make sure that it is performing at least at the same level as your current model, right?
Before Braintrust, we had no specific way of doing it because we didn't have scores set up, we didn't have, you know, sort of like the evaluation. So, with the usage of Braintrust, what it has enabled us to do is to simulate how the performance of the lower model would look like and we've used Braintrust extensively to run offline evaluation and see what the effect is going to be and also online evaluation to see that the intended effect that we observed in offline is is as expected.
So, that's I think one of the use cases. The other one is more generic, which is like it would have taken us like a lot long to evaluate a new feature shipping into Travel Assistant, but with Braintrust, we have been able to make sure that we are assured that the new experience for the users going to be good. So, Braintrust has helped us a lot in shipping fast. >> Cool. And also, the other part is observability for sure.
We do We do use Braintrust for observability if this one is working. Yes. So, yeah, this is true example from from Braintrust. We are spoiling basically the workshop with whatever you are going to see there. But, yeah, we can we can track everything regarding tool calling, the other agent, the in terms of number, in terms of quality. That Those kind of insights you will need them to get from proof of concept to actual system in production or in your company or for something actually that's been used there out there and users find it I would say true working product and not I would say just any uh a uh AI-generated system.
No, we are we need we'll need that for definitely for for those cases. Uh would you like to add something here? >> I was just going to say that BrainTrust enables you to look inside complex agentic workflows up to like tool call level, token level, which is like very insightful and helps you debug a lot of things in production and before you deploy it. >> Yeah. And one last thing probably is the cross-functional uh uh friendly uh point.
Uh that's something that we have discovered along the way uh because we are building those systems from travel assistant to the models. At some point, we need to have we are a big company and we have uh people from product, people with non-technical backgrounds, and we need to communicate and we need to share many things and they also need to self-serve some of those things. We cannot also say babysit any uh people and say, "Hey, you should do this.
You should do this. Let me get you the logs, download it, and send them that." That does not work. At scale, uh we need like a way like where we can help us say uh can work cross-functional way um and let people be I would say um uh uh free to do whatever they want and self-serve with with data and insights. So, this is what we have also discovered uh along the way and BrainTrust helps us I would say uh with many requests that we that we had, but yeah, really appreciate it uh for sure.
And of course, more we are using just a part of uh uh that system. We have our own things. Uh BrainTrust, I think uh you are building even more uh more more toolings. So, so definitely uh more to come. I think the workshop will definitely help you discover all those uh things. And uh of course, have fun building uh during the workshop. I'll say laptops out and if you are interested, Trainline is hiring uh in the AI engineering side, of course.
And uh give it to you, Shah. >> Perfect. Well, thank you so much for that, team. Yeah. >> And just to say again, on behalf of thank you uh and extension of my wife, we thank you so much for the split saver functionality. Thank you so uh give me kudos to that. All right then, team. All right, let's proceed with the the setup. So, again, this is going to be a hands-on workshop. Uh better dust off your bash skills, team.
Jokes aside, I've got a nice raffle commons to help out with that. So, hopefully everybody, or at least most folks, have joined the Slack uh organization for AI engineer and then also join the the channel. I've put a link to this, uh but there's also a QR code to the repository which we'll be using. So, it's there. So, I'll give it a few minutes. Um the really key things is um signing up to a free Brain Trust uh account.
Um if you use Gmail, you can use the little plus sign trick to just create a new account. If you um you know, keep it If you are an existing Brain Trust uh user today, uh so hopefully it doesn't pollute your uh existing one. You will also need access to an OpenAI API key. Um this should be easy for you to generate. If for some reason you're able to generate a key, then please let my colleagues know. We will be able to send it and DM you on on Slack to use for the specific session today.
Um and I think obviously that as an engineering conference, and especially in AI, uh if you are using AI coding assistants, feel free to use that with your IDE or in the terminal to guide you along the way, ask questions on the code base, and so forth. Um I've tried to simplify a lot of the scaffolding in place. So, I'm using Mise to uh manage uh Node uh and PNPM um on on the machine. But, alternatively, you can download uh Node v 20.2 and the specific version there.
It should be fine, uh but again, I just want to caveat, especially for this workshop, I fixed that specific uh let's say runtime. I'm using make to uh just against syntactic sugar to wrap the commands. If for reason you don't have make installed if you're on a Windows machine, you can then just run the PM PM commands which are then just wrappers for the package.json file. So yeah, we'll give it a few minutes for folks to to scan and clone.
I just want to point out as well I have a in the sheet there I can see lots of folks have already gone through it but yeah, we provide a step-by-step walk through. So as mentioned as we're proceeding with the workshop today if you are stuck, you can please ask questions on the Slack channel but also you refer to this. As mentioned this will stay this is a public facing asset so even if you outside of this workshop you're stuck, you can kind of go to it at your own leisure.
Just kind of maybe show of hands who here is not able to get an open API key? Open AI sorry API key. Hopefully folks will be able to provision that. It's really critical to this. Yep. >> There are some people who joined late who don't have the QR code. >> For the Slack? >> For the Slack yeah. >> I can do that yeah. Okay folks I think we did have a few late starters so yeah, I didn't want to spend too long with this but if you did join late, please scan that QR code uh, which will give you access to the AI Slack engineering Uh uh, organization and then join the AI channel uh, AI engineer Europe 2026 BrainTrust workshop.
It's a public-facing channel. Through that channel, you'll be able to get access to the pinned um, uh, content. So, the repository as well as a cheat sheet which will be following as well. Just yeah, just a bit of a sense check folks. Um, hopefully most of you would been able to join the Slack channel and were able to clone the repository, have access to that. All right, then. Um, as I mentioned at the start of the um, the agenda as well, what we'll be doing is creating a um, an example, let's say support triage agent.
So, you know, hopefully this kind of example-wise a lot of uh, systems that folks may be have already built or or trying to build. Um, it is it is a fictitious uh, application uh, designed specifically for this workshop. Please do not use it in production. It's just really around to just teach us um, you know, how how do we build these AI patterns in for scale. So, the idea there is um, given uh, a ticket might be raised in the particular system.
Uh, we have a set of agents that goes through a a pipeline process. I'm saying pipeline but a stage process tool calling uh, that produces um, a set of information that can be emitted into downstream systems uh, from that perspective. Okay. Um, just to kind of help visualize what uh, the system does. Again, mentioning we're taking the ticket input. Uh, we have a first step which is collecting the context. It's quite a deterministic way of extracting information.
We then proceed with the agentic type of portion where we've broken down into three stages. So, we have an LLM and two calls to to triage a particular issue. Um, we then want to can a policy review to make sure the output is is correct. We then want to create and draft, let's say a customer-facing reply with reply writer. And then we want to package this together and depending on how severe this particular ticket, we can invoke another tool call to whether we need to escalate this to, let's say, a human in the loop and then draft the final result.
So, again, not too too complex, but just give you an idea of of the the system that we'll be trying to build and and operationalize in the session. One thing to point out is, you know, as we begin to use BrainTrust, it's like really where does it fit into helping you, you know, deploy and manage this at scale? I just isolated this here where, you know, everything we're going to be doing towards the end, especially as we progress, we'll trace it into end as my colleagues have talked about.
We'll be able to use management for prompts, so we're offloading that what you'll do locally into a secure environment. Tool calls will will be managed. So, again, how do you, you know, talk about getting to external systems? And lastly, around the evaluations and scores, which will be run through our BrainTrust infrastructure. It's worth pointing out as well, if you do go into the GitHub repository, there's a page GitHub pages version of this as well.
So, again, the slides will be readily available for you to use and consume. Okay. Key thing as well with this is I'm trying to help you um do each phase and checkpoint this. So, if you ever stuck, the idea is just use get get checkout to a specific tag or the branch. But in this case, each individual tag or branch is fully runnable at that stage. So, if you ever stuck, again, check get checkout to that branch, you you make install, make setup, run the commands, and then you should be able to get near identical output every every every time.
So, again, this is all documented in the the read me, but it's also included in the cheat sheet as well for you to progress. Okay. So, yeah, just to kind of give you the sequence of events, uh we'll be doing kind of a scaffold and and set up. Uh I'll talk about building a basic agent. Again, this workshop isn't really designed to talk about building agents. It's about, okay, once you do have an agent, how do you then operationalize that?
So, but just for brevity, we want to kind of do it step-by-step. And then we'll kind of get into the the core fit where we'll add a tracing, talk about the evaluations, you know, talking about the the golden set and identifying where we can can improve. And then we tie it all together and using kind of manage infrastructure within Braintrust to to help operationalize this out. And again, talk about that collaboration that my colleagues at Train Line talked about.
Also, kind of the key thing which is like once you do identify a production failure, it's then how do we apply a fix to this and then complete that flywheel and give you a finalized asset. Okay. So, in step one, we're going to be talking about building the agent. So, in this checkpoint here, um again, we need to start somewhere, right? So, what do we do? Uh we will take an initial call to an LLM, we'll set up a prompt, one shot, one in, one out, and get an output.
Uh so, again, if you're doing a proof of concept, this might be great as part of that initial spike, but again, we know there's there's work to be done. But, yeah, for context, we're just going to be stepping through this. But, as I mentioned, just because it works in the demo, doesn't mean it's necessarily going to work in production. Um I'll put some pseudo code here just to articulate here, but you can see here it's uh as as with all uh these um you know, language models, we provide a a set of a system prompt, we provide uh user, and especially the text, and we pass that output back into our application.
So, one function, one model call, uh we have the structured output which we want to receive as part of the the output. Okay. So, what we'll do as well when we do the the scaffolding and checking out to the first branch, uh we can run a set of people tickets. So, it's it's a case that's kind of already done in code or you can use the command line. So, I'll just demonstrate that shortly what it would look like uh with the outputs.
And you'll get something similar to that in JSON format. Okay. Oop. So, hopefully everyone can see this at the moment. Um I have the the application here. Okay. So, I'm just going to go to basic checkout. Okay. And uh what's this what is mode for? So, if I take a look here, so I'm now checked out into uh that first hike, so building the basic agent. Um I've got the application here, which is again very similar to the the prompt code.
Um we're creating the client uh calling open AI SDK. Again, for brevity, I've I've just I didn't use any agent agent SDKs, but we fully support that as well as we proceed with the the workshop. Um So, this is really available and you'll see again from make file uh if you don't have make installed, then you can just execute uh PNPM scripts or I don't know if you if you use NPM or Yarn, that's also possible as well. I haven't tested it, but theoretically it should work cuz it's just using package.json.
So, in this case, let's say you want to do something like, you know, make ticket. So, when I give it a thing, say, you know, my password needs to be reset. Um in this case, I've provided some defaults, so I'll just enter, enter, enter. And now it's just making a call to open AI. I'm just using um, you probably would have seen from the environment variable file, I'm using GPT 5 Mini. You can switch it if you wanted to, but uh, just for the purpose of this, we'll keep it simple.
Um, so you can see here, I've got the ticket and it's provided some some output here as well. But again, that's a singular shot. I think we've run out of time. Okay. So, uh, you know, based on this output, you know, it looks fairly plausible, uh, but again, it's not going to account for a lot of the edge cases that we want, especially if you've got a lot of nuance to the organization, which we're trying to build into uh, the logic here.
Uh, I can even do things like make uh, demo. So, make demo. Go to scripts. Um, yeah, it's the same thing. I'm just calling the same function, but I've got it codified as JSON here. So, yeah, the JSON fields are available to see. Awesome. Okay, so that's fairly straightforward. Uh, don't want to dwell on that too much. So, the next thing I want to do then is talk about um, adding kind of local tools. So, we probably want to say, "Look, let's try to make this a bit more deterministic." Even though our prompt might be very well structured, I may you to bring in uh, different ways of how it might operate.
So, uh in this case, uh I'm calling three different tools to look at relevant um help desk articles around that. Uh this could be both internal and external. I mean, we look at certain things that have happened to the account. So, let's say a customer might have done a certain migration and that might have an impact on back-end systems. And there's probably reason why I want to, you know, create an escalation here. And in the case of what I've done is I've made it a bit more deterministic.
For the purposes of this workshop, again, this is kind of treated as code, but uh in reality, you are probably are going to be interfacing with external systems like the vector search, MCP CLI, and and other types of uh uh interfaces to to build out uh this capability. And again, key thing here is uh the more things that you add, the number of ways that it can fail will also increase. So, again, this is why tracing, as we uh begin to go, will become more important.
Right. So, just one here. So, again, read me here. Next thing we'll do is go to our local tools. So, what do I check out. All right. So, in this case, yeah, the tools that I created uh are then available here. But again, they just kind of checked in as code for simplicity sake. And again, I can do the same thing where you you make uh ticket. Say, password needs resetting. Account locked. >> Okay, now see it's provided a little bit more information.
You can see it's a little bit more verbose because we've introduced tool calls into this and it's giving more context to the the LLM. Also worth pointing if you are feeling stuck folks that a lot of these what all the the tags of the workshop branches are built sequentially. So if you let's go into let's say get tag number six, it's going to include everything that's part of that. So don't feel like you have to go through each one.
If you are feeling stuck and you want to skip, you can do that as well. Okay. Let's get on to tools. So again, we already I've already showed the code anyway. So but this is an idea to pseudo code what it would look like. Um stages. So I think this is the next thing where again we're breaking down that monolithic call for an LLM. We're now introducing tools and the next thing is even further drilling down to special stages of how the LLM should behave.
So you would have seen from that sequence diagram where done effectively five stages for this where I'm now setting up things to collect the context, triage, determining if it's meeting brand policies, providing a customer friendly reply, and also something internally for our systems, and then finalizing the result for downstream systems. Um again, it's just coming from you know traditional software engineering. You're breaking down your problem.
You can see the exactly where something is going wrong in the stack and you'll be able to remediate. So, again, where possible, try to be more explicit and break it down into challenges that you can work on. So, yeah, just a bit of pseudo code. Um This is kind of like what it would look like. Um probably going to put a debugger um with a new IDE and and take a look at that. Okay. So, I've kind of been doing this as in part, but uh We've already started with the um start of the point and then I'm going to talk about, you know, doing the the specialist stage here.
So, let's take a look at the next part of the read me. So, get check out specialist stages here. So, if I take a look now, um in the source folder, it's it's got the individual uh functions which are are pieced out. Uh the prompts associated with that uh is being used. So, let's say here, it's a triage from what we're using. And if I go down to the application, you can see here um uh it's it's down here. So, I'm using a asynchronous functions to uh execute.
Yeah. And similarly as well, if I do something like, you know, make ticket. Maybe I'll do something a bit different. What's another classical problem? Uh let's just say, I need to up grade grade my plan from pro to enterprise, but the website is not working. Getting 500 errors. And you know, so in this case, I'm in customer tier two. Uh we're talking about um billing here. And my account is actually account number three.
So, again, just to show you that we are This is live. It's not doing something that's hardcoded in the system. Um, worth pointing out as well, this stage is slightly slow because we've broken down the individual one-shot LLM into sequential calls. So, it is expected to take a little bit longer, but again, that's just part of building out this agentic flow. And now you can see again, it's a bit more verbose, but it's There's a lot more thought into, um, this this agent here.
So, I mean, as you can see again, we we talked about a billing issue. Um, you can see that it believes that it's quite high because a customer wants to do that upgrade from a different tier, but there's obviously an impact to this from a revenue perspective. So, in this case, we should escalate to the appropriate, uh, people on our side. So, you can see here what the escalation reason should be. Um, and there both the internal and the and the customer are facing reply as well.
So, we've included a confidence score, um, to say, you know, if this is true issue. But again, um, as mentioned, um, tool calls will be able to pull information which is happening across different systems to provide a a greater level of confidence. Okay. Perfect. All right. With that in mind, um, so, you know, again, we've now hopefully shown how we can build and take an agent, break it down, build something that's multi-stage and introducing tool calling to give us what we need.
The next thing is then to provide that information and start tracing it. So, we can actually see what's happening, you know, down to the individual details. And this is where, you know, observability piece comes into play. So, what we want to do in this kind of section is break down the full execution path. We know that the stuff is very nested in structure. So, there's two calls, there's additional function calls behind it.
We do want to track some of the key things and I think, you know, Myun pointed out, you know, the early struggles that they had to talk about, you know, latent latency, cost, tokens account, especially for time to first token is a very important metrics that we see many of our customers trying to identify. Also, you know, what were the inputs, what were the outputs, metadata associated with this. And then also including additional types of fields so that, again, when we talk about monitoring and observability, we can query it within the UI on the fly and serialize it if needed.
So, again, having an output is not enough. We need to understand the full execution path and that's what tracing allows us to do. Yeah, just to give you an idea of the context and I know it's a very contentious topic topic to some of because again, I come from an a background in full stack development. So, you know, I both coins and tools both, you know, Python Go as well as TypeScript, but yeah, we our SDKs is multilingual.
So, Ruby, Go, I think even dot net. We've got some folks who are using it and that's that's all good. But yeah, I've just kept TypeScript for simplicity here today, but yes, our SDKs do cover a range of different languages out there that you can start tracing your application with. So, yeah, Uh real key thing to this and uh one of my challenges I do see with our customers working broadly is they will do an individual interaction uh against a a singular uh parent span.
And this even might work with say multi-turn, multi-conversational agents. What we want to be able to do is trace that into a nested structure so you can see everything but in one interaction, let's say one conversation uh in a in a parent span. So, it's really critical as we start instrumenting our application to make sure that we're getting the the right structures in place. Otherwise, you're not you're not going to be able to see the full effect of where things might go wrong in your application.
Okay. Um again, this is just come down to reading the trace. Hopefully, folks um especially I know some of the you've got more software engineers in the room uh that are using kind of traditional observability tools out there, this is not too dissimilar. Uh but for folks who maybe new to this uh screen, a trace just allows us to uh not find out what has happened but what's currently happening with the application uh in in in real time.
So, uh again, tracking every single call, bringing that metadata and again, depending where the failure mode is, we can then identify, okay, what might be the the course of remediation that we need to take. Okay. So, in the next stage, what we're going to do is add tracing to this. Um and then we're actually going to run it and then go into Brain Trust and hopefully, we'll see it happening. But something I do want to bear in mind before I do that.
So, So, let's go to our read me. It's telling us to uh tracing. I think it's fine. Okay, now we're good. And all right. >> Okay, so I've introduced a helper script here called tracing. So um what we do quite well within our SDK is again we want to not introduce more complexity where it's needed. So if you are using again the standard um LLM providers SDK, you can simply just wrap that function. We provide that out of the box.
But then when it comes down to the individual uh calls again we can set up uh some some helper uh scripts here to help and just kind of wrap up uh as needed. If you are using something like Python, then you can also have a a decorator function which helps as well. In this case, you know uh I've got a nice little uh function which helps with parent and child. And then when it comes down to the actual um application tracing uh you can see here like I've got a um a child span which has been um executed throughout this.
Uh Again, I don't want to do too many code demos at this point because the code code is uh expanding pretty heavily. Uh but yeah, just kind of want to walk through the kind of key concepts for this. Uh Okay. Um before I run the uh application, I do want to go over into the the Braintrust UI. Sometimes it might create a uh if you've already signed up for a free trial account, so you would have done that earlier today.
It'll create like a test project. That's totally fine. We'll create a new project as we execute this. Uh really key thing uh as well um if you haven't done already top left hand corner you go to your profile um Oh. Okay, let me just create a new one. So okay. Let's create that just for now. Uh go to your profile and then where it says um, API keys. um You know, enter your uh a name for the key, generate that in, and then use that within your environment, uh your your .env file.
Or secure that uh in a key vault if you have that already. There's also the OpenAI key which we're hoping you generated. Um, that should also be used uh in AI providers. So, I've set this at the organization level. You can also do this per project, uh you know, depending if you want to segregate it by particular team or environment. Uh that's also fully supported. Uh but in this case, um we'll need this key for later when it comes down to the managed online scoring.
So, but just a bit of a tidbit, um that key that you generated, make sure you put it into your uh BrainTrust um organization here to be used later. Right. With that in mind, I'm going to run the demo. So, here and so. And uh just point of reference, if you are ever stuck with subs, you should just use make setup, make sure everything's in place. In this case, then we want to do kind of make demo. I just we're going to run and execute those tickets.
I can even do it uh using the uh make ticket command as well. Okay. So, while it's running in the background, oops, sorry. While it's running in the background, um if you go back to uh your application, you would have see uh there's a project called helper workshop. So, that's the one that's will be created as part of this uh workshop here today. Um if If navigate to the logs tab, you'll start to see this coming through in real time.
So, what pointing out again, thanks to capabilities of BrainTrust, we have near instantaneous right into our system and read available shortly after. So, especially it's it's a non-blocking function. So, a lot of our customers, especially the more sophisticated ones, are really using BrainTrust at scale and really pushing the envelope. So, it's not the case I have I've got a thousand traces, you know, they push you know, tens of millions of traces at a at a time across a very short period and they want to be able to get an aggregate this.
And again, as you build more sophistication in your application, you're sending this out to more users, that's going to grow up pretty quickly and you need to be able to have a system that can handle this at scale. So, once you start to see the logs that are coming in, come in reverse order, I'm just going to take a bit look at the first one here and we'll start to see the instrumentate application we've already traced.
Um There's a particular button here that allows you to view this in full screen, which I think is quite helpful. So, as I mentioned, because we've got this nested structure in place where we're going through each in the sort of tree, this interaction at the top is the is the is the demo ticket. You can see at the very top level we're saying how long it took to actually run that invocation, the prompt tokens in, out, cost and latency associated with that as well.
I've also included some metadata because I want to be able to kind of extract and filter that out as needed and I'll show you how that work. Um and everything is available here to view. Um metadata is here. If I also want to take a different look at this, we can even look at individual steps. In this case, I want to look at what's happened with uh tree of specialist down to the actual um to the LLM. So, again, SDK uh provides a lot of um flexibility around this, so you can see what was put in.
Uh the reasoning behind it, and the information that we set up on. And then, coming down to the last, you know, should we escalate or not? Um another view that we will see as part of this is uh taking a look at the timeline. So, just gives you an idea like a waterfall methodology to see if there's any particular step which is taking longer, do we need to remediate? Uh it's impossible to do. Um All right. So, if I take a look at my logs page, let's get about out of that.
Refresh. So, I'm thinking about kind of four tickets that were pushed. Yep, they should come through right now. Okay. Yep, so that's It's two tickets at the moment, so again, that same information that we we played is also available in the console. Okay. Yep, so we've covered tracing. Let's talk about the evaluation portion for this. Okay. It's quite interesting uh again, depending where you are in your journey of building this application, uh a lot of customers already build an application already, have it sort of monitored and pull that in.
But, what happens when you have uh effectively a cold start problem where you don't know what you're building, right? So, what's what does good look like? Um and effectively, what does good good look like to me? Um in the case of the support application that we're developing today, um kind of can non-negotiables, you know, have we categorized the support case? Have we made sure that there's no severity uh low severity outcome for uh these issues which are blocking.
Um does the escalation stay in policy with with SLAs? Um does the structure look uh look sound? And if we're making any particular changes, does it actually improve, uh, the way we would actually breaking or having a regression uh, to to the application. Um, we can do this using evaluation. So, an evaluation, um, for those hopefully many folks will know what they are, but those folks in in the room, think of it as a way you have your your data set, your input, you have a task, and then you have an outcome which you want to evaluate against a or scoring function.
Uh, and this is kind of a little bit different comparing say traditional software development uh, into working with AI systems because of its non-deterministic nature. So, in this particular, uh, portion, what I'm going to talk about here is creating what we call a golden data set. So, in this support application, I've been kind of testing it anecdotally, but I want to create a a set of edge cases where I think this is really going to help give us at at least an initial level of confidence to the business that what we're releasing out into production, you know, is is sort of fit for purpose.
We're always there's always room for improvement, but I'm going to be functional I want to just think, look, it's not just me releasing this application based on vibes. I have a concrete way of saying, okay, this is how it's is performed over time. Um, to do this, again, we use kind of two main types of scoring functions. The first of which is deterministic. So, I think a lot of folks have may have already started using this.
Again, coming from traditional software engineering, uh, you know, unit tests, uh, I wouldn't say analogous, but they are quite similar in nature where, um, they're very easy to run, uh, cost effective. You're not actually using a model at this point. The secondary type, which is a little bit more sophisticated, which is then using an LLM as a judge, so another AI system. And this is really helpful when it comes to systems where there's nuance which can't really be determined on deterministic systems alone.
So, again, creating one to talk about, you know, branding style is just meeting, you know, customer satisfaction uh and so forth. So, the most important thing is if you cannot write it in a deterministic way, you want to be using LLM as a judge where possible. So, and why this is more important, again, is just making sure that any change that we make is a safety change as we progress this. Okay. So, I'm going to pivot into the idea again.
Let's take a look at the uh readme file. And then we're going to do Oh, that's a dictionary maybe coming from. I'll see. Let's check out. There we go. Okay. And then based on this here, we'll make demo make setup. We should be fine. Okay. So, one thing you I'd like you to do as well is then to run the seed dataset command. So, we do make seed dataset. Okay. So, what this is going to upload our evaluation uh test cases into into the Vercel UI.
So, if I pivot into uh my datasets now, you'll see it's called helper seed dataset. Again, just for simplicity, I've created 10 um inputs. Um I've also categorized them uh around there. So, you input what we expect and some metadata associated with that. So, kind of the core structure to uh creating an evaluation. >> So, um as part of this as well, um and the deterministic and non-deterministic, I've created some scoring functions, uh which we'll be using as part of this.
So, it's available here. Uh and then the app scoring functions. So, again, checking the category uh as a schema in place. Um Is an escalation reason when when needed? Again, very easy to run uh and codify. Uh yeah, we have a little more sophistication with the the customer rubric. So, this is more of the LLM as a judge use case here. All right. Uh then what we can do is then go back to the readme. So, we don't need to do the demo.
We already pushed that through. Uh let's go ahead and then do um make eval. Okay. So, this is then going to run an evaluation. Uh let me say here, um the seed data as well. Uh again, you don't necessarily have to put it into code, but um whether it's coming from a database or in this case, I've just used a flat file JSON for this. Uh just to give you some context. So, we're up running this evaluation. >> And we should have if we go to UI a place for experiments.
So I'll just double check here. Yeah, so that experiment ran against all the JSONs associated with it. So we should you should get a an output like this at least in the terminal. But if we go back to the UI uh again you can we're starting to track the our application against the inputs and outputs for this particular data set. Um quite similar to the um the tracing you you would have you would have seen from the online uh traces.
You get to see a very similar view here uh and see how it works across um your experimentation as well. Yeah, and I think yeah, pointing out as we progress through the workshop uh we'll start to see how do we then improve it using the the sort of different functions and track that across the UI. If you need a bit more real estate as well, um you can collapse the menu uh which does help. Oop. Okay, time. Perfect. Okay.
Um All right, let's proceed with the next sec- checkpoint uh around deploying and managing this. So, um as I mentioned, I think a lot of folks, at least anecdotally when I when I've seen this from the customers, is again, things will work really well on my machine. Okay, now I'm checking that into code. I kind of want to take this into a place where um I can start to collaborate a little bit better. It's in a There's a point of reference.
Um I've got versioning history. I can start to identify, again, who's made what change. And you need a way to be able to bring these users together. So, I think as Salma talked about our collaboration. Um what's interesting as well is, again, changing the prompt on your machine and then trying to ship that code to a repository and there may be somebody who's, let's say, a non-technical SME, a product manager, perhaps.
They want to update the prompts. They can't do it. They have to tap you on the shoulder. Maybe that's happened to some of the folks in the room here today. I know it's happened to me a few times before. And I can I can I can certainly can get really frustrating. So, we actually just want a way to be able to pull that together. Really key thing, especially for those who work in very regulated industries, like reproducibility is a very big key thing.
And I've worked in, better part of over a decade in in both banking and and capital markets. I know obviously with uh you know, the regulatory out there, especially uh things like right to be forgotten, understanding, you know, who's made a change, especially in a stressed exit scenario, how can we put this into a new system? This is really key to to helping unlock that. And interesting when it comes to to identifying changes before you do that.
So, again, we just want to We don't want to just be making changes, pushing out production, and asking us what's happened. We we need to be able to kind of get that in place. Um and so for us, again, uh we're introducing some capabilities now. So, what you've been running at the moment, you've been running the tools, you've been running the prompts, everything on your local machine. What we want to do is then offload that capability into BrainTrust.
So, when your application is running in a secure environment, it can refer to BrainTrust to uh pull that information uh and then uh help with the path of execution, which again follows the tracing mechanisms that we've done. So, um by default, when you're running the make commands, um the runtime mode will set to kind of local. If you want to use managed mode, just use the prefix managed and whatever the remaining make commands uh will happen there as well.
So, what I want to do then is um pivot into um the IDE. So, going back to make uh my readme. Okay. I'll slow I'm not taking I'll see. Okay. So, git checkout Oh. Okay. So, just take a look here. We'll do make setup. All right. And the key thing I do want to emphasize at this point because we're managing this in in BrainTrust, use the setup um BrainTrust um command here. So, what that's going to do is going to package up those scoring functions, those tools, those prompts, and push it onto the secured uh infrastructure.
So, you should get an output like this. And what that would look like uh in the UI. So, I'll just go back to overview to see. So, on the left-hand side, um if we take a look at our prompts, you'll start to see those three prompts that we created as as part of that that workflow. Um, give you an idea here. If we go to the, uh, the tree yard specialist, so, um, you can define a slug. So, let's say a an immutable ID, which you can then refer to it in the code.
Um, this can also be generated if needed. Um, the prompt, um, is again treated as code. Uh, we also use some interpolation there if you want to parameterize this. I'll show what that looks like shortly. If I go to the scoring function, uh, so let's click on scores. Uh, okay, now that will pop come up in a second. Um, take a look at parameters. So, I'm just going to take a look here and I've created a parameter specifically, just to simplify things, changing the baseline model.
So, maybe just kind of a show of hands, um, you know, these models get released so quickly. Has anybody had like a PM or let's say, uh, let's say a non-engineering SME say, "Hey, by the way, can I change uh, the prompt Can I change a model or the prompt and see what that would look like? Has that kind of happened to you folks?" I think we've got a few hands here. Great. Okay. So, the good thing with this now using the the managed parameters, those non- uh, technical SMEs, they can come into BrainTrust and change the prompt here, write a comment to say, "Look, um, let's use, um, a different model.
Let me use something like 4 mini." Uh, just say, "Testing a new model." I'm going to save this version here. I'll write the comment. And what I'm going to do as well, just to give you an ending thing, I'm going to do is, you know, make set of BrainTrust. So, every time you change something, um, you can run this command, but I'm just doing it for brevity here. Um, that's more to keep, um, this model in sync. So, if I go to prompts, uh, you know, uh, that's kind of the sync, but if I take a look at, um, the, uh, uh, sorry, the parameters in place.
Yeah, it should be there. And, so, sorry, all I want to do is find do something like, um, uh, I think this is what was our runtime model. Go to unlock. So, in this case, by running the map manage mode, I'm just changing the course of the execution, uh, not to run the model locally, but to follow the path uh, of what Brainwave Trust has, uh, set for with the model. Okay. And, uh, if everything goes well, take take two minutes just now.
You can see that I've changed the model here. So, again, I didn't have to do any code changes. All I had to do was go into UI, change the model I wanted to, any other parameter, run that, and have that, uh, use as a uh, establishing baseline for the evaluations. >> Um again, if you want to uh you can also run the the demo script to push in the demo tickets. I'm just skipping that for the workshop today. Um I think I do see this with many of our different customers is saying, "Oh, you know, um there could be not necessarily a cause of concern.
It's saying, 'Hey, by the way, Braintrust is now having access control to these putting parameters.'" Um what I would say is Braintrust is not really intended to replace that rigor. You probably still want to use things like version control systems anyway to track that. What we're just saying is when it comes to operationalize it and making sure that other users are able to work on a shared system, this is a recommended path that we would take uh to help out with that.
So, again, you would probably still have to have your prompts, your tooling, your parameters in a kind of centralized way, but then provide automation in place for you to synchronize that and work. And that's the best way we've seen customers take advantage of this. Okay. Uh next portion we're going to talk about online scoring. So, now that we have those evaluations in in place, um what we're going to do is then apply those um scoring to actual live production logs that are coming through uh in the application.
So, okay, it's great that we've done our test cases. We We've got some level of confidence that it's working. But again, there's no substitute for production data, right? We all know this. So, what we're going to be doing is uh creating uh again, moving that logic uh into Braintrust and setting up what we call automations that would then track and uh evaluate this as logs are coming in in in real time. Uh worth pointing out, when you start your journey, um it's probably, especially if you're using uh LLM as a judge, you want to start with the let's say a higher sampling rate.
So, again, as logs are coming in, uh, where possible you you want to make sure you identify a baseline. But, again, there's a trade-off when these costs can be quite expensive, especially if you're using more sophisticated models and you need a higher rate rate of reasoning. At that point, when you do want to, uh, are happy with, um, the output, you really want to reduce the the sampling rate down to 5 to 10%. So, again, you you're managing your your costs effectively.
Uh, deterministic scores, again, they're cheap. Recommend running them all all the time. S- Sorry, what was that question? Yes, so I'll bring that up in the UI. I'll show you. Yeah. Okay. So, if I go back into the, uh, IDE, I'll go into readme here. Um, ah, so that's why I missed out on manage tools. So, let's just say I get check out. Okay, there. And as I mentioned, they all build up on each other, so skipping this is totally fine.
So, get check out. All right, now, if I do, um, little make setup, everything should be fine. Make uh, setup. Bring cost. Um, So, in this case, I'm actually taking the the tools as well, um, for production. Again, this is just really help accelerate things where possible. Um, coming down to, um, if I hit refresh, um, right. So, our scores, the scoring functions that you saw earlier, uh, in code, then I'm managing BrainTrust.
So, you can see it's it's available here. Um the ones that I want to call out is the the triage. So, that I've got an LLM as a judge, um which has been applied. So, putting an output taking this. And there's an automation rule in place. So, root quality online. Um so, what I've done here again I've automated the setup as as code. So, to to bootstrap the project. But again, uh what we're going to do here is to say, look, depending if it's an individual span, we can run the execution or the entire trace.
And this is why metadata is so important because we we may only want to trace maybe specific failures that might happen within the code. Uh again, depending on the use case. Um my sampling rate, as I mentioned, is set to to 100. But again, for more expensive calls, we want to take that up as well. So, this is what the automation uh or how we would do that in in BrainTrust today. >> Click packaging the scores and >> No, so automation is more like the execution against uh incoming logs.
So, it's um with that. Yeah. But I'm I'm more more holistic I've applied automation to setting up this environment and scaffolding. Yeah. So, that that's just want to delineate that, yeah. So, there's a question here. >> Yeah, what kind of things are you scoring? >> Um could you expand on that, please? Yeah. >> So, the automations are like all of this judge, right? >> Yes, correct. In >> If there's like no ground truth to it, >> Okay. >> like how can you perform online if you don't have ground truth? >> Yeah.
Well, we probably want to take that as an edge case, push that into a data set, identify it, and then move that back. So, that would be the approach that I would take for this. Um if you don't have any ground truth or data already, right? So, this is why I said, depending where you start, it's better to have some kind of data, and then begin your flywheel around from that. >> How do I suppose if you have some data, how do you apply that to the analyst judge decisions? >> Oh, that case.
Um we can probably put that into a data set, and then replay that through the playground. That's That's going to be the way we would do that. Yeah. Yeah. All right. Um so, we checked out online scoring, coherence over time. Okay. Now, to the remediation portion. So, hopefully seeing the delta or again, why we we had it today. So, hopefully this might help you uh with that particular question. So, again, here's something that might happen as a as a plausible input to our agentic system is, you know, customer uh user might say, "Hey, no, this isn't urgent, but I'll see if I can't export the invoices um before uh board of meeting." The model says, "Look, hey, this looks This looks okay for me." Someone says, "Not urgent." Come si, come sa.
But, the business is very different, right? This probably does need uh immediate attention. Uh your CFO, I'm sure there's an end of quarter report that needs to be done. And this is the difference between uh what we're doing here today is trying to identify um what is a proper failure mode, and then remediate that where possible. Um so, in this case, again, I can run um um this particular mode here. I've got this in in this in this data set.
So, uh just to kind of give you a a play of scene where we want to replay the failure, we want to specific evaluation against this. We want to tighten the prompt, run it again, and see uh what that looks like. We probably want to use it against not just a one particular test case, but across our entire test cases as well just to see if if it doesn't work as intended and we have regressed on something else from that perspective. >> Okay, so let's go ahead and have two We have two separate branches for this.
So, I'll split out branches A. One's going to have the the failure rate in mind. We're going to go to the UI and view that. Um and then we're going to talk about the the remediation part uh there as well. Okay. Awesome. I think So, just go over here. So, many production videos. See. Put here. Okay. So, um mix up. Oh. Okay. So, we can even do runtime mode. So, runtime mode. merged And then we'll uh replay failure. Okay.
So, I've got this set of five cases which are, you know, the regression of failure modes in this case. So, that first one that you saw in the example ticket, that's this done as a JSON file here. So, if I take a look here, okay, so if I take a look at the the failure mode here, I can drill into this. CD And you'll notice as well that that we set up the online the managed tools as well as the online scoring, the trace becomes even more sophisticated in the fact that we're executing this against the secure a brain trust environment.
So, again, moving from local to to managed. All right. I'm just going to go here to Uh, the make file. Sorry. package.json Let's go to Here is another file. to my knowledge. It's another file here. home loading home impact Let's run the evaluation here. So, I'm just doing an evaluation against a specific um scenario. Is that right? >> Do you also want to show the monitor? It's about latency and things like that. >> We'll do that. >> Just >> Okay.
Okay. Um yep. So, as we can see as we progressing with uh the application, the experiment uh again, viewer just allows us to see the progress of our changes. Um so, you can see we've run the the latest set. We notice some, you know, degradation here, which allows us to track it better. So, the managed data set that I had in mind, uh the various failure rates have been captured, and you can see, you know, I have the ability to compare it against existing uh experiments to see to track the progress and remediate where possible.
Okay. So, forgive me, but the next thing is I'm going to go to the readme, and then proceed with the uh remediation. Okay. So, do git checkout -b issue next up. All right. Just Okay. So, in this remediation, I've changed um the the prompt which I've used. Uh So, uh one way to to view that is if I take a take a look at the the prompt and the the change. So, let's see. So, you'll see any any differences there. Just replace.
Okay. Come down to the uh remediation file here. Let's say I do one say say collects here. Let's give this a say say. on the the trust command This collects. There's a specific flag. I think I don't know if you want to just double-check with that. Yeah, that's it. If If it fits So, I don't know if you want to just put that if if exists replace um in the end of the chat. Yeah. Yeah. Yeah, no, no, it's fine. So, yeah, just focus on If you want to in the part of the remediation script, have you set the environment variable BrainTrust underscore if underscore exists, you set that to replace, um it's going to push in the the updated changes into the into the environment.
So, then if I take Oh, sorry for that. Okay. Um and then you'll see as well, now that we've updated the prompt for your code and pushed it up, uh it's um it's available here as well. Um so, as well in the UI, um again part of the operation operationalization, uh we can see you know, who's changed what, but actually what has been changed to the particular prompt as in this in this case. Take a look here. So, we're including two facts.
Clean that out. Okay. So, we're coming back to here. Let's say we didn't do a um We're on this evaluation again. So, we made the change to the prompt. And now we're running um through remediated version to see how that performs. >> Perfect. Okay, so that's run the experiment here with the new changes, hopefully. And I'm pivoting to the the UI. I'm going to go to experiments here. And it's barely there, but you can just see, you know, kind of taper back up.
Just, you know, any improvement going up is an improvement. But yeah, just to give you an idea here that um you know, now that we've done the evaluation, uh actually what I can do is um do a diff. Uh Um we can do some do a comparison in the delta. Yeah. Yeah. So, yeah, that's come up uh improve over time. Um uh which we're which we're uh the intended outcome. Yeah. Okay. So, I I think we're approaching the end of the the content.
Um so, I know it's uh been a number of steps, so I really want to thank you for your time and attention to kind of walk through that. Uh as mentioned, the the artifacts are public. We've got the cheat sheet there. We've got the Slack channel to help if you have any any questions, but uh hopefully just to give you a summary of what you've uh accomplished today uh in in this order is you went from taking a single-shot prompt into building a five-stage uh AI agentic workflow using uh tool calls.
What we're able to then do is then inspect how this works pretty much end-to-end diving into this by adding Braintrust tracing, making sure everything is recorded. Um we also then want to talk about, you know, how do we then evaluate the system from our um you know, when there's no online when there's something new, uh, creating those those uh effectively golden set with those test cases but you going to execute against.
We then deployed, uh, that uh, those managed prompts, those tools and parameters into the BrainTrust secure architecture uh, to able to use and we've also added online scoring to then evaluate the system uh, as it unfolds. And then we picked a particular uh, production failure. We looked at the trace. Uh, we modified the prompt in our uh, case and we saw the delta there. And running the evaluation again we saw it achieve uh, back up to to where it needed to be in this case completing the full evaluation of your building, observing it, deploying it and taking note of that moving forward.
Um, so yeah, just to kind of uh, call it uh, and bring it home. So again, hopefully this is not uncommon but again, what might work in production is not really going to work in prototype. Uh, we really need to break this down, identify the failure modes and and move forward and that's where again explicit stages become really really important, right? And what um, again, this does introduce more uh, areas of of where things could go wrong but it's easier to debug if that's the case.
Um, yeah, there's no substitute for diving into the code and tracking everything. So I would say it's it's observability is table stakes at this point. If you've got a production in application and you're not tracing it, you need to go back to the drawing board and get that done operational and hopefully again we can show you how to to be able to do that um, using BrainTrust. Um, again, no substitution for production logs but better to start somewhere from nowhere.
If you have an idea of what an issue might happen, these are your perfect ways to to start an evaluation. So using those uh, failure modes as as your test cases and as I mentioned this is a continuous process, right? Nothing is ever done. If you've ever worked in a job development, constant feedbacks is is important and again, we're bringing this this operation model but with uh a new surface of of of operating. Um and yeah, just hopefully bring it home to your teams here today.
So, um you know, my encouragement to you is to if this is something of of of interest is pick something that's already operational that it doesn't have to be the entire suite. Maybe start off with something that's maybe a bit more um uh that I guess more critical that you really want to improve um the operational modes. Uh the tracing, collect your edge cases from that that mode, um build the automatic scores, and then boot everything back as possible.
So, again, the faster feedback loop that you have, uh you have more insight into the more that you can again improve the overall uh delivery and operations of of your system. Um yeah, and just kind of call to action. So, again, I know we've thrown a lot of content at you. It's uh uh we'll obviously try to get a bit of feedback. We'll try to put some tables in place. So, I appreciate everyone's kind of juggling everything to do this, but um I guess as mentioned that you have to start somewhere.
Let's just try to accelerate you. Uh we have a list of documentation that's provided. Um that's uh you can also use our AI agent on the agent to to search if you have any questions, but we also have a cookbook available. So, I tend to throw that cookbook directly into, you know, cut Cosa Codex or whatever and say, "Hey, based on this, take the SDK, uh and start tracing my application. It does a pretty effective We even do have uh version announced a CLI um for our Brain Trust um application.
So, that allows you to even do things as auto instrumentation. I'm just going to plug my colleague Eric, who's doing some fantastic work, and please check out his booth cuz it's uh it's it's amazing. Um and again, if you this is something that's interested you and you want to explore more, then you know, please reach out to your account team at Brain Trust. Again, we're happy here to support where where possible. Uh and if you're on Discord, you're feel free to to join.
I'm happy to answer questions there from that. And yeah, again, we really want to thank you for your time, attention, and energy. I know it's a really sunny day and I I don't want to keep people in here. I want you to get some fresh air, but yeah, just on behalf of Brain Trust, Train Line, and myself, like thank you so much for your time and attention. It's It's been an honor and look forward to seeing you out there tracing and and getting value from delivering AI in production.
Thank you.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.