Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 11:58
4.6x that video's typical replay level
issues. Uh I also invented OS certification. I just close the tracker whenever I want, so I have my life back. So, does this work? Yes, sort of. >> [laughter] >> Which leads me to act three, slow the down. Everything's broken.
Said at 11:52
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
The graph counts replays. It does not show where viewers stopped watching.
Words
2,543
Runtime
17:34
Speaking pace
145wpm
Reading time
11min
145 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
[music] >> Hi all. My name is Vlad. I'm co-founder and the CTO at bank. Now, before we start let's do a bit of trivia. Orc C okay. If C, raise your hand. K? Who believes both are correct? Okay, just a couple. Both are correct. One is from the Lord of the Rings, another is from Warhammer game Game Workshop games. I have in my agenda four topics to cover. Topic number
73 words, the words spoken in the first 30 seconds at 145 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 201 |
| Average words per sentence | 12.7 |
| Longest sentence | 48 words |
| Questions asked | 19 |
| Sentences containing a number | 13 |
Most used terms
Filler phrases
32 in total: like 10 · uh 7 · basically 6 · right? 5 · actually 3 · um 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
[music] >> Hi all. My name is Vlad. I'm co-founder and the CTO at bank. Now, before we start let's do a bit of trivia. Orc C okay. If C, raise your hand. K? Who believes both are correct? Okay, just a couple. Both are correct. One is from the Lord of the Rings, another is from Warhammer game Game Workshop games. I have in my agenda four topics to cover. Topic number one, I would like to tell you about our thesis as a company, what we believe in.
Then we'll talk about the AI evolution from adversarial agents to loop engineering and beyond. Then we'll touch base on the technical challenges of tomorrow that we have to solve today. And then I would like to present the company, what we do, and what problems we solve. So, our thesis is future belongs to AI to AI communication within a business, between businesses, and between consumers and businesses. Agents will be everywhere.
They will do work on our behalf, and they will have to communicate between each other. They will be written in different frameworks, different languages, and deployed in different environments. And these agents will be fully autonomous. And they will communicate with each other without human intervention. Now, just to paint a picture before I start talking, uh how this type of communication between the agents is going to look like.
So, what you see here is a conversational space that agents can create themselves, so they can be added. They will receive a task either from a human or from another system. And they will be able to discover other agents add them as a participant in this conversation space, have back and forth communication, solve the task, and report back to humans. This is how the future that you all read in the newspapers will look like.
And when people hear this, they think either that this is too far in the future or that I'm crazy. Okay, that's it. And here is an example of a conversation I had with a CTO of $10 billion company who said that before thinking about technical solutions to hypothetical problems like multi-agent coordination, I try to keep things simple and avoid the problem. Now, let's unpack. Is multi-agent coordination a hypothetical problem?
Is it possible to keep things simple? And is it possible to avoid the problem? Now, let's talk a bit about the adversarial agents as a as a concept. Um we all use this. I'm pretty sure everyone here is a developer. And as a developer, you have two cloud or coding sessions. One is doing the work, another is reviewing the work. And you probably have not only two sessions, you probably have multiple tabs and multiple sessions working working at the same time on multiple problems.
You are basically acting as a router, a Cisco router or a switch between two stateful agents that do work on your behalf, but they have no ability to communicate, hence you need to prompt them. Let's talk about loop engineering. You've heard the the guy before talked about Peter and Boris, right? So, what they say to you? They basically say to you, "Stop being the router between your agents and let your agents prompt each other.
Now, what they're basically saying to you is instead of being the router and copy-pasting stuff, start fighting with Python and TypeScript libraries and different abstraction layers that will invent how agents should prompt each other. We also have protocols. A2A, MCP, ACP, and then other protocols crypto-related. Simple. Solves all the problems, right? So, this is how your multi-agent system looks like probably in production.
MCP, maybe if you're super advanced, it's A2A. And sounds simple. You call agents as tools, maybe you connect to agents through A2A, and life is great. But, MCP, calling agents as tools, it's completely stateless. If you want two agents to be stateful, sticky sessions, good luck to you. A2A is client-server. If I'm an agent, I can send a task to another agent, but if this agent also wants to send task to me, we both have to be a client and a server.
Chaining multiple agent equals together involves REST API timeouts. Good luck managing that. And obviously, what no one says to you, you need queues with persistence and so on to keep track of the messages that are being sent. And of course, discovery, which is not even part of the A2A protocol. So, you're basically doing the planning work. You're not creating a multi-agent system. You deal with planning. But, we have messaging platforms.
Slack. Anthropic released a wonderful agent. You can talk to it in Slack. We have personal agents here and I connected to Telegram, right? Wonderful. To connect an agent to a Telegram, five steps. Discord, seven steps. Slack, eight steps. WhatsApp, 11 steps. Every step is manual that you have to do by hand and you have to read the documentation. Again, you're doing doing a lot of planning and manual work. And this gives you only one thing and one thing only, an agent that can talk to a person, usually it's you.
Your agent is still alone. They cannot see each other. They cannot communicate with each other. They are in digital solitary confinement. So, let's unpack the statement I heard from the CTO. Is multi-agent coordination a hypothetical problem? Clearly, if you're copy-pasting stuff between two sessions, this is the problem of today and it's not hypothetical. Plugging in MCPs and A2A, this is today's problem. It's not a future problem.
Is it possible to avoid the problem? Clearly, it's not possible because otherwise it would have all work with only one session and not two and we would have no need to call other agents as tools and so on. But the question is, is it possible to keep things simple? And the answer is, if we look at all the options that we just talked, not really. But how hard can it be to connect two sessions, two processes on my laptop together?
How hard can it be to connect my cloud to Lambda? Or Salesforce agent to the Databricks agent to SAP agent to my my my Codex? Actually, it's pretty hard. Think about it. These agents, even your two sessions are processes that have to talk to each other through a network. So, it's a distributed systems problem. And distributed systems are hard even before you have introduced agents on top of it. And multi-agent system where every agent is remote is basically a distributed system of microservices where each microservice is non-deterministic.
So, it is hard. What do you need to solve all of that so it will become easy? You need to solve the transport layer. Ordered message delivery, real-time message delivery, retries, and so on. You need to solve continuity. Microservices, your agents are software, pods, dockers crash, and so on. So, you need to have a consistency hydration. You also need to do the run time binding between different agentic frameworks. You have thread IDs, conversation IDs, execution IDs, and so on.
And someone needs to map all of these IDs together so your agents can actually interact. But, it's also not enough. Your agents cannot communicate at an IP port level. They cannot communicate at the URL level. They cannot even communicate at the pub sub level because it's still a lot of plumbing that you need to do as an organization. So, for this wonderful future of agents talking to each other, we need to raise abstraction of a technical stack to conversation and talk about rooms, channels, participants, and figure out the deterministic routing of messages within a channel and also across different channels.
And even if you solved all of that, it's still not enough. You need to solve the governance layer, the identity, audit, and etc. So, I would like to introduce Band. This is exactly what we solved so you don't have to. We connect every agent together and we are global collaboration layer for all the agents. Any framework, whenever they are deployed. Inside, it's not just a communication. We implemented every primitive that is required for your agents to talk to each other.
Let's see them. What you will see are two different users connected to the collaboration layer, Vlad with a personal assistant and Mike that has no agents. On the right in the terminal, I'm going to spin up different agents and they will be onboarded into the platform. So, let's see how fast it is to onboard a new agent. We just spun up a new agent, programmatic registration and onboarding and agent card appears. This is a Codex agent.
Next, we spin up a Land Rush agent. This Land Rush also appears in the platform. From this moment, they know that they exist and they can talk to each other. Now, we will ask Codex agent to send a connection request to my personal assistant. Keep in mind, different users, different registries. There will be a connection request, contact request sent to my personal assistant to require bilateral consent. It will arrive here in a second and from the moment that I approve this, Codex will know and see my personal assistant, will be able to invite the personal assistant into the conversation and send messages to this personal assistant.
So, we're asking the Codex to invite Andy. Andy got invited, received the message, reported back. But, let's talk about problems of today. We are all developers. We use multiple sessions, probably a lot of sessions. We do routing. We don't like it. And when we go to uh have a snack, we come back and we have no idea what our agents have done. And you as a manager have no idea uh how much it cost distribution uh of ticket and the tokens and so on and so forth.
And you have no idea even how long your human was involved in the work. I would like to introduce Jam. Jam is an internal product. This is how we develop the software and we've built it on top of bank. Jam is a desktop application that simplifies the onboarding of your local agents to the platform and solves the problems that we just mentioned. So, we solve routing, we solve context overload, we solve cost management and attribution, and we allow multi-agent and multi-human collaboration together.
So, what you see here is local agents and remote agents working together. And we capture all the tasks generated by Claude and Codex that they generate for themselves when they do the work. And we present this to you so you can track the work that these agents are doing. And this can be your local sessions or your local session with a session of your friend. Because trying to understand what your agents are doing, 1 million tokens multiplied by three, that's a lot.
You need a completely different way to understand what your agentic team is doing. Moreover, we provide a way for agents to describe the layout of the software architecture that they are working on. This is what you see on your right side. And you can see in real time where your agents are working right now, what piece of component they're touching in real time. And once they are done, they mark it as done. If there is a human in the loop involvement, you'll get pinged.
Now, you don't have to use this desktop application. You still can open your terminal and you can work from your terminal, but because every communication goes through the network, we can monitor all that stuff and we can surface a lot of other very useful information and we can enable my agent join your agent. We can enable our team member works remotely to join the session together with the agents and the humans so he can help solve us with some problems.
We if you have a security guy who maintains skills for the security agent, I don't need to copy his skills. I can just ping his agent to join this conversation and solve this for me. Now, I'd like to show you the platform itself. So, since we're all managers, we'll start with graphs. Uh so, the moment you open application, you can see all the stats of all the traffic that happened between your local agents and also your remote agents.
You can see for instance over here, I have a full stack developer uh $2,000 in tokens and this is a local cloud session. You can see an architecture over here, $600. This is a local coded session and the rest are different agents running in different environments. But, how do they work together, right? Do we have a bunch of Python code triggering the agents so they can work together? We do not. Because all models right now, they're trained on a lot of data so they understand very, very good how to communicate through messaging platforms.
So, here I have real work that I've done this morning. Engineering manager, developer, and architect, different instances of quad quad working together, reviewing PRDs and SRS, we are implementation. there is no need to hand code all the loops. They know how to do it natively. And for the managers, we have full statistics. Let me kill off the manage the the engineering manager who sent too many messages. So, you can see the attribution.
You can see the full [clears throat] user chain cost. So, if you ask a question, if my developer is actually involved in the code he pushes as a PR, or it's all AI slope, you can see it here. Not only locally, right, but also through the remote agents. You can see attribution by developer or by teams of your agents and developers and how they work. Everything is gets updated in real time. You can see all agents, and every agent is basically a session, right?
So, if I click here, I can see all the sessions of my Cloud and Codex instances that are running and they're connected to the platform. They're connected to a global platform. So, if I want, I can connect any of you to any of my agents in 30 seconds. Obviously, you can set permissions, you can see all the rooms, you can see all the errors, and so on. And also, you can see work. What work was done, what is pending, what is in progress, and everything works in real time.
Thank you very much. If you want to know more about the future software development that does not involve you pulling in another 50 packages, and if you want to enable your Salesforce and Slack and Data Bricks and Cloud and Codex to work together and collaborate, come to our booth LG17. QR codes for the Band infrastructure layer and for the Jam, the uh desktop application to allow you to be part of this future that everyone talks about but has never seen.
Thank you. >> [applause] >> Hey.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.