Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AWS Developers · @awsdevelopers
Words
5,138
Runtime
46:23
Speaking pace
111wpm
Reading time
21min
111 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Every time your AI agent responds, you are paying for the words going in and the words coming out. On your bill, you will see those called tokens. The more tokens you send in, the more you pay. And if you what you send in is not quite right: too much or missing something important, your
56 words, the words spoken in the first 30 seconds at 111 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 408 |
| Average words per sentence | 12.6 |
| Longest sentence | 68 words |
| Questions asked | 12 |
| Sentences containing a number | 38 |
Most used terms
Filler phrases
14 in total: like 7 · you know 3 · actually 2 · uh 2.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, published by the channel, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
Every time your AI agent responds, you are paying for the words going in and the words coming out. On your bill, you will see those called tokens. The more tokens you send in, the more you pay. And if you what you send in is not quite right: too much or missing something important, your agent starts to hallucinate. There are five techniques to help you to reduce token waste, improve accuracy and catch failure before user see them.
And each one is a code change, not a prompt change . Let's see. The first one is semantic tool selection. You filter which tools going into the context on every call. The model only sees what it needs for that specific query. Two, Graph-RAG for precise query like aggregation, counts, multi-hop reasoning, you replace text retrieval with structured graph queries and the model gets a computed verifiable answer. Not a sample.
Three:Multi-Agent Validation. A second agent checks every response before it reaches the users. Silence failure get caught instead of confirmed. Four: Neurosymbolic guardians. Your rules live in Python code. Not in the prompt. The model cannot skip. Five runtime guardrails. Steer don't block. When a rule fires, the agent self corrects and completes the task. No hard stop, no user retries. For each technique. I will show you the agent without it and then with it so we can compare.
The demos use a travel agent build with a Strands Agent (of course).. An open source agent framework from AWS. In this demo, I'm using OpenAI as the model provider, but you can swap it for another supported provider with one line change. All the code is open source in the repo linked below and at the end I will show you Amazon AgentCore AWS managed service for running AI agents at scale. I walk through the architecture and show you where to find everything in the repo if you want to deploy it yourself.
Before we start, if you have not seen the previous video where Morgan talks about Amazon Bedrock and Strands Agent from a scratch, I recommend you start there. This builds directly on the top of it and of course, the link is in the description. Let's get into it. Our travel agent has 31 tools. Flights, hotels, payment, weather, cancellations. Every time a user send a message , all 31 tool descriptions go into the context window.
The model reads all of them before deciding what to do. And if your agent has memory that grows too. Every conversation, adds more context that gets sent with every single message and you pay for every single one of those tokens whether the model ends up using that tool or not. To understand where those tokens come from, you need to see what a tool actually looks like to the model. In Strands Agents, you write a function with the tool decorator and name a description in the docstring, typed parameters.
Strands takes that and generates a schema: name, description, parameters - and that schema is what goes into the context window on every call. Each tool schema is about 70 to 100 tokens depending on how many parameters it has. If our travel agent has 31 tools, that adds to up to somewhere, I don't know, 3,000 tokens per call just for the tool description before your messenger, before the response, every single call. By creating a tool database, we can filter the tools that the agent may need before the agent is invoked.
With the filtering, the model sees only the three more relevant tools. Token usage drops from thousands to fewer than 300. Let me show you in the code. Before running, open the requirements file. This demo needs Strands Agents with open AI sentence transformers, the local embedding model we use to cover text to vector. This is open source so you can run that in your computer. We use fast CPU, the local index for similarity search, and python-dotenv for the API key.
Install everything with only one command. Let's do that. Let's come back to the Jupyter notebook. Here you can see that I'm using OpenAI as model provider just right here. To do this, I need to import the OpenAI library just right here. And Strands uses Amazon Bedrock by default. If you do not specify a model, it picks up your AWS credential and uses Claude via Amazon Bedrock. And if you want to use another model provider, for example, Anthropic or Llama locally, you need to install the requirement for that model provider as we did before for OpenAI.
And you need to import the library for that model provider. In this specific demo, I'm using GPT4 mini. And if you want to runBedrock , you only need to delete this and delete the import library. Two lines change. Everything else stay the same. Now here in this application I have the register app. In register app is where the filter layer life. We have built an index that takes your list of tools. For each one, it combines the name and docstring into a text string and converts the string to a vector using sentence transformers.
The model I just showed you before that you can run locally as open source and those vectors or tools vectors go into the files index creating our vector store of tools and we use search tool that takes the user's messenger the user query to convert it query a vector with the same model. So we can use the query to search the closest tool in the index and this application returns the actual function object ready to pass directly into the agent and we use swap tools to make this work across a conversation.
This function helps us to add and remove the tools in the agent and the conversation history stay intact. Only the available tools change using the swap tools. Let's come back to the Jupyter notebook. Here I have three tests on the same query and we are using this question, this test query just right here. They are like the ground through. We know the query and we know the the tools that are going to answer that query because we need to we need to know what is the accuracy of this agent.
So I run this. I have this helper function. And let's try it. The first test is with my traditional approach with all the tools. So I going to send all the tools and for this demo I am using 29 tools. When you run this you can see the query and the amount of tokens that this agent is using to answer that question. Here you can see that all of those are between 2,000 and 1,000 tokens in each call. Now let's run the same but this time with a semantic approach.
If you compare the first question, search for something in Paris, the agent is only using 300 or almost 400 tokens for this invocation. And the same question in the other agent, the agent is using 2,000 tokens. So yes, we have a difference here and the accuracy is a lot different as well. For the first one the agent have 31% of accuracy and for the second one the agent is having yeah a little more. All this invocation we are using this only we send a question the agent respond and then we are using another invocation to the next question.
What happen if we want to have memory? Yes, we use the swap tool application. Let me show you. So here, every time that I search a tool, I do the same, receive the tools and I swap the tools. So every time that I invoke this agent, I'm sending all the context and the chat history. And you can see here in this demo that amounts of token is every time more because you are not going to send only the tools, you are sending the chats history for the old conversation and that's it.
In production with Amazon Bedrock AgentCore you don't need to build this index yourself this is only a demo that runs in my computer. You can use AgentCore Gateway the routing layer inside AgentCore. It handles tool selection automatically. You register your tools once and it finds the right one for each request. Same principle, no infrastructure to manage. RAG Retrieval Augmented Generation is how agents access your own data, your knowledge base.
You take the user's question, search your document for the most similar content using vector search and pass what you find to the model. The model answers for that. It works well for open questions. Find me something about this topic. But there is a category of question where RAG breaks down. What is the average rating across all hotels in Paris? How many hotels have a pool? Vector search always return something even when nothing is truly relevant and the agent only sees the top N chunks for all your data at time.
It can not aggregate, count or traverse relationships across the full data set. So it estimates and it present that estimate as a fact with full confidence. Here, let's see a comparison. Left side RAG retrieves three chunks from 300 documents and the model guesses. Right side Graph-RAG runs a query across all the data and returns a computed result. Graph-RAG addresses this differently. Instead of retrieving text chunks, small little pieces of your data, you can build a knowledge graph from the documents.
Nodes, relationships, structured data. And for this demo, I use Neo4j, locally, and the model. My model, my tool writes a Cypher query to search in my graph database. Cypher is Neo4j's query language. It's similar to SQL. Yeah, similar but it's not the same. And the graph runs the query across all the data without using tokens. and without using the model and then the model gets back a computed verified result. Not a sample.
Let me show you that. First you need to install the requirement we use again a Strands agent with OpenAI, neo4j, neo 4j Graph-RAG, faiss-cpu because we are going to compare two different agents. One that use vector store and vector store is going to be faiss and we use again sentence transformer to do the vector embedding. For my demo, I already have the graph built for this demo. But in the code I share with you how you can build it yourself.
Let me show you. You can use build_graph_lite for like a for a little version that only use 30 documents to build the graph knowledge base. And you can use the normal version of the hub to build the database with a graph base with the 300 documents. Let me show you. The repo has step by step to set up everything for you and this is specific to neo4j. Remember this only work with that and to build this knowledge graph neo4j allows you to use a llm that can understand your data and it create the nodes for you.
So you don't need to hardcode the node for this graph database. This is only specific to neo4j. The llm here in this demo is OpenAI. But this library also supports anthropic mistral and of course Amazon Bedrock. The key thing here remember you don't need to define the schema because the llms retest and discover the entities and relationship by itself. Let's test it. I have two agents same model same data. Bedrock agent has a tool that searches by similarity and returns the closest chunks.
The Graph-RAG agent has a tool that runs cypher query against neo4j. Same model, different tools, different data access. Let's see. Here is my file data store. I have three 300 documents for that. And for Neo4j I have uh 296 hotels in my knowledge graph. So I have my tools the one that allows my agent to search hotels using vector similarity and the other one capable to create and run the cypher queries. I run this and first test aggregation average guest rating across all hotels in Paris.
What is the average guest rating across hotels in Paris? That's my question. So I run this. I have my Strands agent for RAG and my Strands agent for graph agent. So let's see my first agent it says the average rating for hotel in Paris basing on the available data are are as follows. So this agent used the word Paris to search in my document data in my document vector store. is received to different hotels but the thing here it's my LLM my agent is the one who is doing the calculation here and this what happen if I have more than two or three hotels the agent is not going to use the other hotels to do this calculation and he's doing all the calculation to calculate the average guest rating blah blah blah blah blah blah.
So it generate the question. For the knowledge graph, I only have the answer because the agent already had the average ready to use. The agent don't need to think mathematical things and the agent is doing this calculation -not the agent - the tool is doing this calculation using all the data in my notes. Now the second one versus counting. How many hotels have a swimming pool? How many? Let's run our test. For my traditional RAG, I received the available data did not mention any hotels in Paris with swimming pool amenities.
So blah blah blah. So it gave me a lot of words there. I only ask you how many. If you don't have the data you don't need to generate tokens there and the other one told me the Graph-RAG there are no hotels that offer a swimming pool as amenity I have the question twice but that's it the graph it go into the tools it do the cypher query and the answer was zero. No swimming pools here. So that's it. The agent don't need to give me a nothing answers.
Now test three out of domain detection. Let's see what has happened here. Tell me about hotels in Antarctica. A spoiler. There are no hotels in Antarctica. There are zero hotels there. Let's see what has happened. So let's run let's fight agents for the best answer. So a traditional RAG agent says it seemed I didn't find any specific information hotel or tell for accommodation option in Antarctica in the available data but it say accommodation options in Antarctica.
You know they don't have data but it recommend me to search for research station, expedition cruise luxury lodge.. come on that no exist there. you know uh it's hallucinating an answer for me. That's a normal behavior for these types of agents. So it start to hallucinate answer when the RAG the similarity search didn't give you a answer that can use to answer a question. Let's see what is happened with my Graph-RAG agent.
So it runs the query and it received zero. There are still no hotel listing in Antarctica. It appears that there are maybe limited or no accommodation. No accommodation officially register in the region. And the other one it give me some recommendation to do my search. So that's it. You can run this code. You can compare. You can use another data set to run your test. Everything is there. So when to use Graph-RAG and when to stick with regular RAG.
Please use Graph-Rag when you want to count, average, multi-hop reasoning. Anything that requires a precise answer and when to stick with RAG. Come on it's not so bad. when you want to answer open question, fuzzy search and a stricter test. But the best system is both. Sometimes an agent fails and nobody finds out. It calls a tool, the tool returns an error, the agent does not surface that error and it generate a confident success response instead.
The user thinks it's worked but it did not work. The agent adds and validates its own output in the same loop. There is no separation no second option. So when something goes wrong it rationalizes and tells you it worked. Here is what's happening inside a single agent where it fails. It calls the tool, gets the error, rationalizes and returns a success response. The user never sees the error. You can address this is by adding a validation layer. multi-agent validation.
Three agents in sequence, one acts, one checks, one approves or rejects. Strands has a built in class for this called Swarm. It manages the hand off between agents automatically. Just define the role of each agent in its system prompt. Let me show you how. Let's go to the Jupyter notebook. For this demo, you only need the Strands Agent for OpenAI. So, you don't need to install another one. So, I going to run my OpenAI key.
The key important here is the Swarm from a Strands agent multi- agent library. This is what lets you connect multiple a in a chain and manage handle between them automatically. Without it, you will have to wire the aent together manually. black. Let me show you how you can create your swam agentic application. In this demo, we are going to use three different agents. [clears throat] The executor who is going to call the tools and process an initial response.
The second one it's the validator who is going to validate the answer that execution just received. For example, is the right tool? I receive the right answer for that tool. And the last one is the critic. Who is going to approve or reject the answer? And to create or swam, we only need this line. We swam or three engines, the executor, palidator, and critic. And we can add a mass hand off. So in this set, in this time we are going to use six hand off.
So the swams handless the flow between them. You don't need to to do nothing else. Now for this fight agent fight we need to compare a single agent we or multi- aent application for that we are going to use four different scenarios. The first is val booking B grand hotel for Ellis for two nights. Then we are going to try to book the luxury resort for both for Bob for one night. And then we want to book the R Paris for Carlo for 39.
And the last one is missing booking log up. And we are going to send the query get details for booking. This is a a scenario we ground true. So for the first one is true and false for the other one. So we are going to compare the result between the single agent and our multi- aent. Let's see what is happen in the comparation. Single agent versus multi- aent swap. You run the code and let's see the fight. So you can see that for the single agent everything is correct but we know that that's not the right answer and for the multi- aent swamp we have some clear verict some hallucination are reject and one missing the issue.
What we have here the executor gets the error the validator catch it and the critic rejected. So the user never sees a fabric response with a single agent we have like a no entry and return success. So every time is going to be success for the swam executioner get the error validator says hey come on hall elucination and critic reject the user sees a clear failure no fabrication confirmation when you use multi- aent validation that catch fabrication after they happen you have a rule for your agent let's say maximum 10 gas per reservation.
You write it in the system promp. You even write it in the tool description and the agent still calls the tool with 50. This is no because it is ignoring you is because prompts are suggestions. No constraints. The model process them as text. No as a logic it has the to execute only code executes logic a rule in the prom the model reads it as suggestion the rule in the code the model cannot escape it neurosyolic guardians put the rules in code agent has a special future call it hooks functions that trans automatically at a specific moment in the end loop.
In this case, right before a tool execute. So you write a rule, check the parameters and if it fails, you cancel the call. Let me show you what that looks like in the code. The key important In this demo are these three things. We have the hook provider is the strand base class you extend to create a hook. We have the hook register. It's what strands passes you to register your call back. and before tool call event. It's the event that fires every time the model is about to execute a tool.
The last one, it's what make it possible to intercept the call before it runs. Let's compare left rule in the prompt. The model read list it as text. It may follow it or not. Ry ruling the code. Python executes it. The model can no skip it. Each rule it's a data class with a name, a condition function and a message. The condition takes a context dictionary and returns true or false. Poor Python code, no model code. Let's see the alien here.
I go to my Jupyter notebook again. I'm going to show you with you everything so you can run it by yourself. So let's see the agent to create the agent with hooks. First you need to define the symbolic rule, the rule that is going to be in the code. So you need that you do that with this. We have our rules. Let me show you the rules. These are the rule that the LLM is going to take to give us a decision. So we have the condition function and we have the rules.
For example, name validates condition validate checks. Validate checks code. You must have check in, check out and message me. Check in must be before check out. Yeah, right. And let's come back. We run this. Then you create a validation hook using the register hook using the before to call even just right here using the rules that I just show you. Then we define the clean tools for booking hotel process payments configure booking everything is there.
So that's are the tools that the agent is going to try to invoke and step four we create the agents for for commission you know because this is a fight between a agent with hook and a agent without and this agent is going to fight using different scenarios. So we have the scenario who is going to configure booking without payment book hotel existing guest limit but bookings and each scenario it have the answer we should expect let's run the first test confirm it booking without payment confir booking blah blah blah so has been success execute without hook hit confir the booking.
But what happened when I have the hook in my head? It seemed that payment for booking needs to verify before the confirmation can be processed. Will you like to process the payment? If so, blah blah blah. Of course, I'm not going to booking something if you didn't pay. So, the code is actually there. Book hotel exceeding guest limit. The limit is 10. So let's run this. I want to book a hotel for 50 people buying the date.
The Grand Hotel has been successful booking for 50 guests. Come on, man. I just told you that the max of our rooms are 10. So it's another error here. And when I using my hooks, the agent said it's appear that the grand hotel has maximum limited for 10 guests per booking and booking must be made or less one day in advance. Oh yeah, that's another thing that I put in there in the real in the rules. When you are not using hooks, all rules pass. probably it's going to block something if they use the problem in the right way.
In this demo just use the same model, same tools and the same pro. I didn't change the prom for my agents. The different outcomes is because the rules are in the Python code in the code in the hook. knowing the problem. This part of enforcing rules in code before tools run and it's also what Amazon agent core policy service does at the infrastructure level. The same concept but managed for you in production. Hooks are all or nothing. they block but something you want to a to adjust and keep going no stop and leave the user waiting that is what I going to show you in the next chapter hooks block unconditionally the agent stop and the user has to try for a hard constraint that is exactly what you And but sometimes the rule is softer.
Maybe a room fits for for gas, but a group of this could book two rooms or a fly is full but there is capacity on the next one. You don't want to block everything. You want the alien to find a option and complete the task that is a steering. Here on the left the hook fires and the task fail it block. On the right agent control steer the model and the task complete. The other different is operational with hooks. Changing a rule means changing code and redeploying with a control or steering which is the name of the open source library using here.
Rules are registered on a local server via API. You only have to update then without touching the a core the end picks them up almost immediately. Let me show you that in the code. This demo needs one extra packet the end core SDK for trans engine the framework that we are using here. And you need to run the setup controls once before the demo to register the steering rules on the local server because we are going to create a local server here in my computer.
In the setup control, you can see all the rules that this engine is going to use to do the steering. So we import the alien control client, the aliens control and in here the control rules. So the name is still mass gas for example. So the definition is guide agent to reduce gas count when exceeding maximum of 10. We have another one that it say deny no payment. So block booking confirmation without pure payment. That's a hard rule.
And yeah, we can put hard rule here in a steering or steering configuration here in my Jupyter notebook where I create my agent. I have two key importance for the agent control SDK. I have the a control plugin that captures a events and sends then to the a control server or rules. And we have the a control steering handle that is listened for a steering decision for the server and delivers them back to the model. together they are what connect a trans agent to the agent control stealing logic.
Now let's see when we compare both approach side by side. The hook version with mass gas hook that cancel the call if we have more than 10 and the aent control version with the same aent same tool but two plugins added. One captures events, one deliver the steering message about the model and the steering rule it is registered on the agent control server. So let's test with my aen with hooks. We have the mass guest hook that block everything else that excess maximum of 10.
So I run this and I going to send the query and the query is boot ground hotel for 50 gas. The engine round the hooks and it block the calls because the maximum is 10. So with engine control you remember that it can steer if I have more than 10 people that want to use the room. So let's see. We have our libraries here. We have the a we need to initialize our a control server. We have the plugin. We have the steering and that's how we create our agent with a steering rules.
So let's run this. How do you take these two productions without managing service without building infrastructure? That is what I will show you now. Everything we build in the previous demos runs locally for the production version. Amazon betro a core give you a runtime a gateway shortterm and longterm memory and cloudatch observability built in. No servers to manage here. This is the code for all the demos I just show you.
And you will find the architecture. The trans agents run inside the runtime. The gate with roots tools calls to lambda function automatically. The lambda functions are going to be our tools. The steering rules from previous demo live in dynam. So you sh them there and they are live on the next call. No redeploying needed. And for Neo4j you can use Aura DV as an external graph database which also has a free tier. The code is here in this ripple.
You will need AWS credential with Metro access. If you are new to this, start with a notebook. I have a notebook ready for you so you can deploy everything. And if you are comfortable with infrastructure as a code, there is also a CDK stack, a cloud development kit to deploy everything at once. Both options are in the repo. To go deeper on a core, Morgan has a full video link in the description. Don't stop at the demo.
Try going to production with Amazon Petro. Let me bring it all back. Token waste on every request. Wrong tool pics. Fabricate confirmation rules. The model quietly skip it. Hard blocks that stop user instead of helping them. None of those are prom problems. All of them are addressable in code. Five technique beyond the prom. All the code is in rebel linking below. Each demo has a notebook start with demo one through demo five.
Each run about fixing meaning for demoxis you will need AWS credential. Have you tried any of these in your agents? I would love to hear about in the comments. See you in next one and happy building.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.