Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 11:58
4.6x that video's typical replay level
issues. Uh I also invented OS certification. I just close the tracker whenever I want, so I have my life back. So, does this work? Yes, sort of. >> [laughter] >> Which leads me to act three, slow the down. Everything's broken.
Said at 11:52
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
The graph counts replays. It does not show where viewers stopped watching.
Words
3,745
Runtime
20:01
Speaking pace
187wpm
Reading time
16min
187 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
[music] Hello everyone. So my name is Aayush Bhardwaj and I did applied AI for a hedge fund and now I do everything tech plus applied AI for a pharma techch startup because you know the way startups are you have to do everything we have multiple hats. So before I start the session, I would like to do a small survey. Can I get a raise of hands for all the engineers in the room? Okay, that's a tough room. Now can I get a raise of hands for managers? Okay, just to
94 words, the words spoken in the first 30 seconds at 187 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 203 |
| Average words per sentence | 18.4 |
| Longest sentence | 321 words |
| Questions asked | 16 |
| Sentences containing a number | 11 |
Most used terms
Filler phrases
162 in total: like 90 · uh 30 · actually 13 · kind of 11 · I mean 7 · right? 5 · sort of 2 · you know 2 · literally 1 · um 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
[music] Hello everyone. So my name is Aayush Bhardwaj and I did applied AI for a hedge fund and now I do everything tech plus applied AI for a pharma techch startup because you know the way startups are you have to do everything we have multiple hats. So before I start the session, I would like to do a small survey. Can I get a raise of hands for all the engineers in the room? Okay, that's a tough room. Now can I get a raise of hands for managers?
Okay, just to be clear, managing a agent does not count. You have to manage people. Okay, we have few managers as well. Interesting. So this will help me like fine-tune my talk a bit. So today my aim is to take you through the journey of how do you actually build an iterate in applied vertical AI and my experiences from the hedge fund and the pharma techch company. So before delving deep into the recipe, I'll just like take you through what do I even mean by applied vertical AI because I don't know it sounds like a very weird term.
It's like the vertical word is kind of forced. Uh I won't lie, it is. I coined this term probably. So applied AI is like built for so applied vertical is essentially applied AI but built for one very specific industry. It's its aim is to simulate a job of a person in that particular industry in a sense. So an example of applied AI is Google translate which is like general purpose helps you translate. It could be used in education tech and it can have like tons and various sorts of uses.
Whereas Alos which is my employer the pharma tech company we specifically build drugs with AI. So that's a very specific use case. Another examples of applied vertical AI field could be the legal tech firms that are now coming up with you must I'm sure you must have heard about them. So those are like another good examples of applied vertical AI. So when I left the hedge fund right so I was expecting that the world would change for me cuz you know hedge funds are like really fast and really pes sensitive whereas pharma is like okay we're going to take 15 years but we're going to do it right.
Hedge fund was all about like you need to do it fast and mostly right. It does not matter if we lose at one paradigm as long as we are overall winning. Whereas a pharma firm is like we have to be absolutely right. You can take a week more and it was true. It's it's a completely different world. But to your surprise and to mine as well, nothing changed. Actually, my job increased, but the core part of my job applied AI remained the exact same.
And I cannot uh express how surprised I was cuz I thought that it'll be a complete different thing, but apparently it was not. So, uh, so then I spoke to other people as well across legal AI and the people coming up with the prop tech pumps, which is essentially the real estate tech firms. And I realized that everyone is kind of building the applied vertical AI in a very similar way. I could see some steps that could be essentially abstracted out and that's what we'll do today.
So uh before again delving deep into that I received a few reachouts saying are people actually putting agents into production and I was like this is such a wrong question to ask. Everyone is putting agents into production even like 15 year 15 year old kids these days. The question to ask is whether they actually work, whether they actually make or save money, whether they justify their ROI, whether uh they are making way more than the amount we are investing into it like end to end.
And I can say from my anecdotal experience, yes, at the both places I worked, the agent either saved the money or made more money. So with that, let's get started. So the recipe I'll take you through a series of seven steps roughly and try to like make this process as simple as possible. So the first step is formulate the problem. So this is sounds like very trivial but a lot of people specifically startups get this wrong.
They just try to do too much at once. Whereas from what I have learned and what I think a lot of colleagues would agree you need to pick a very narrow task. You just cannot ask it to do everything. A good example for this could be let's say if you build something in finance you won't ask it to like hey can you fetch me top three market opportunities that I could invest in. No that won't work. You have to be like very specific like you pick a market you say let's take the US equities then you pick a industry let's take it and then you ask it to like rank stocks based on some parameters like capital expenditure or let's say the AI um uh investments.
So you pick like very specific things and then you uh sort of formulate a very narrow job for the AI agent to do and you can build like any number of AI agent. Last I checked there was no tax on building more AI agents. So why do you want your single agent to do everything? So this is important and this in the same uh in the pharma context is the exact same. We just break down the process into steps and then ask really pointed questions with the agent.
We model our agent for our task. So once we have our problem right off the way, we know what we're trying to solve. The next step is identify the data. And I cannot stress this enough. This is a really, really, really important step cuz everyone has news data. Everyone has like seller side reports from JP Morgan, Morgan Stanley. Uh, everyone has the AXI preprint server or pupcam or your research papers, right? But what actually makes your application better than let's say chat, GPT or cloud?
It is your proprietary data. So the thing with proprietary data is it's really expensive to buy and most people won't sell it to you. So you need to curate it by yourself. Imagine your organization has been working for 3 years, right? They already have a lot of data. It's just unstructured. And in the age of LLMs, I think this is a very fairly easy task to make unstructured data into structured data. Like a LLM workflow could do it overnight.
So to give you a great example of the proprietary data that finance industry has, it's the trade thesis which is like what trade work and why it worked. And in pharma it is the data for failed experiments because successful experiments data yes you can get it but failed experiments that's relatively hard to get. So now we have the problem we have the data. What's the third step? That is to model the problem like write the prompt.
So while writing prompt you like what we should aim is to model it after the person who you are trying to replace. I mean that's the hypothesis but yeah no offense we're not trying to replace anyone with AI but that's the ideology behind writing prompts encode how a person would solve this job into multiple steps. So it's just like a like a mental model. So this is again fairly simple. Next thing observability I'm sure you have been in this conference at three years and this word I think I don't know you'll be hearing about like a thousandth time there are tons of observability provider if you can see it you can fix it so you need observability to see the traces understand what your uh AI application is doing and debug it so sorry but all of this was the easy part to be honest all of this fits one screen the mythical 10x engineers can do this stuff in minutes like literally this is the code you precisely need to build an AI agent.
So that's why it's not the mode. Uh of course accept your proprietary data. So what do you do now? What do you do after doing the first four steps which is observability and prompts and like uh getting the data right and everything UI trait. Now the thing with iteration is like when I joined the hedge fund I thought how hard it can be. I mean everyone can iterate. I mean we have been iterating our whole life for each of the task.
But to be honest I could build it but I just could not tell if it worked cuz I'm not a trader. I'm not someone who has a PhD in biology or chemistry. I just don't understand what the model is saying. What is the output of my AI agent is? And since most of you are engineers you would relate. You can instantly tell that sonnet 5 sucks because you have your own training. You understand okay this code is not great code. Whereas some X model let's say F five you see okay this is great but not as great as the high pace because you have been trained for this for life.
You have a mental model to judge these things but you just do not have the same kind of mental model when it comes to like predicting trade thesises or doing like really specific task that vertical industry does. And this is also the place where like a lot of vertical AI projects quietly die because on the surface it looks like you have made it, you have built it, let's put this into production and start selling it. But no one would buy it.
The same way you would use an inferior coding model. So as an engineer when I ran into this I just couldn't accept honestly. I thought no there's certainly more that I can do. We don't need other people. So I thought I could uh LLM as a judge my way out of it. >> [sighs and laughter] >> And this was a really really stupid mistake to be honest cuz what LLM is essentially doing it's it's predicting the next probival word.
So if you see it's just like jargoning its way out. It does not understand what alpha means. It does not understand how to actually create value unless you have like taught it somewhere. And whereas a human can just tell it instantly what's and what's not. So I'll just try to delve a bit more deeper on why you can just iterate. So first thing is that model cannot verify itself specifically in these fields because reinforcement learning via verifiable rewards is really good at math and code because you have like answer keys you can verify your code is compiling or not and there's tons of stuff you can just model uh the complete thing around this but when in these fields there is just no way to model it and and let's say if any error gets in it just compounds with every step and that's what licken seems to think as well and Now the more important part that we touched upon previously the data.
So the interesting thing with pharma and finances the data was never there and I'll explain to you why. So any institutional manager holding over $und00 million in qualifying US equities are forced to publicly file their holdings long position holdings every quarter. And once a hedge fund does this, this is the percentage decrease in their returns because everyone just sees those reverse engineers and takes away their mode.
And when it comes to pharma, right? So this is the number of uh so by law you are like required to disclose every clinical trial pass or fail you have done but 30% of the firms which is like nearly one-third of firms never do. And in like 2026, FDA had to like publicly remind over I don't know over 2,000 sponsors that they are I mean doing injustice by not uh releasing unfavorable results because this is the exact data which helps the model thing which helps your LLM actually reason through these complex and niche industries and they hide it because for them it's like a chicken laying golden eggs.
Why would they sell their chicken? So naturally neither open AI nor entropic has that has this data because it's like gate cap. You just cannot hire a trader for $100 an hour and have them annotate that stuff because like lots of NDAs and they definitely earn more. So okay now I have told you about tens of problems right now you naturally think okay yeah right then what do we do? How do we build a startup in like a vertical spare space?
So very self-explanatory. You hire the person who you want to sell it to cuz there is to be honest no other way around. I have tried a lot of stuff. You just need to hire the user. In finance, in a hedge fund, this was very easy because the user was kind of like my boss, the trader. We worked together. But in the pharma tech startup, it was very weird. We are like a bunch of young engineers and we like, "Oh, we need a 20-year-old scientist in our company to tell us what to do." Yeah, I guess we do.
And then we hired someone, right? And that someone actually changed the trajectory of our tools. Our tools started making sense when we pitched to the other pharma companies, the big ones, the big farmer, they started liking our tools because it's kind of spoke their language versus the normal jargon LLM language. So once you have hired the user, let's say then what would you make that user do? You try to build a learning loop out of it.
The domain expert can start at the like a very very low level, the ground level where they just think about prompts. Okay. Yeah. Uh I mean let's not ask LLM to do this. Let's ask a very specific query. Again, they'll help you curate data. Just like engineers know which conference are which are not, which uh research paper sites are great, which are not, which are like top leaders in engineering, which is which are just like influencers.
Similarly, a pharma expert or let's say a trader knows which sources are more reliable than the other. So, they help you create their data. They help you like refine your prompts better and they try to create like thinking models of how they would think about a problem cuz I mean let's say if you if you follow five steps to solve a problem, right? You just cannot do it in in any random order. There has to be a logical flow. there has to be a natural flow that so that's what they uh try to curate like decompose a problem gradually refine and then finally judge so the person who sort of has lived through the complete of the industry that they're trying to revolutionize their judgment is now like turning into agents so that's what's happening behind the loop so uh to do this there are like again multiple ways I mean each of these could have been an hourlong session on its own and I wish I could take but these are like few ways that I identified.
Uh I'll just like take uh you through them like really quickly in the interest of time. So supervised fine-tuning I think most of you would know where like model mimics human demonstrations. Uh reinforcement learning from human feedback is like a kind of uh a very efficient way where human preferences train a reward model. Then rubrics as a reward is I I like to call it reinforcement learning from AI feedback. This is because that you can human can just create a rubric and then AI will just like grade itself based on that rubric and that improve its own processes.
But again there is a slight chance that you might run into an echo chamber with rubrics as rewards and the cheapest of all and I think the highest ROI is the error analysis whereas the observability part that you set up earlier you just analyze the logs plain and simple you understand where model is going wrong and then you just try to correct it. So this is where you have like don't have to touch any weights and the most highest ROI way to get the impact from like start on and once you understand like uh what more you could do or if error analysis is solving or not you can just gradually climb up the ladder and probably uh later on go to the ultimate reinforcement learning from human feedback because that's I think in our industry kind of the golden standard these days that you need to do our LHF to actually get some edge But certainly there are some pitfalls of it like for example now there's GLM 5.2 to right you fine-tuned it right uh Alibaba cloud or let's say Deepseek will release a new model then you have to fine-tune that too as well so there is a cost it's not cheap so once you have done all this you just create a loop and you just like go onto that loop you hired one user you hire more users they ask more queries the scope increases the data increases at this point you are kind of generating your own data the exercise you have been doing in loop right that exercise itself is generating a very I would say a crazy data set of what works and what does not works and this loop never stops once you feel confident enough in your application you just ship it provide it to the external paying users and then you see the magic of it that it actually works so I just pulled this stat from Stanford AI index report because it's a really nice report that gives you an idea of what the state of AI is and this says like 80% 89% of enterprise AI agents never reach production again I disagree every AI reaches production but it just fails to work or like justify its own cost.
So that's the real thing. You can just build and ship AI agents whenever you want but you need to justify ROI and finance and pharma are two such industries where if it does not make money it's shown the door simple they won't like wait and say okay maybe it'll work in two years maybe the cost will be lower in by the third year. No it has to instantly make money. It has to like hit the ground running and if it does not show the door instantly.
So just to summarize the seven steps that I feel are like good enough to give you an abstraction of how the vertical AI industry moves. You formulate the problem statement. You source your data sources. You prompt it well. You refine those prompts. You observe how your tool is performing. You don't iterate yet. You hire the user. And this user or users now play with the tool as much as possible. They like kind of form a learning loop, an endless learning loop that goes on.
And at a point when you feel yeah it's it's really delivering that alpha over let's say claude and chat GPD you just ship it you start earning money. So one more interesting thing. So HITL is like kind of a thing. Everyone is like yeah let's add human in the loop. I would say not yet. Finance and pharma are still those two industries where it's AITL AI in the loop cuz everything is like done by the expert but the AI assistant really helps save time.
Like for example uh it may take an X amount for a trader to form different trade thesises and AI can just give him five candidate trade thesises. But which one would actually work in the market and which won't is the discussion uh the discussion still lies with the trader and same for pharma when you are like picking drug candidates which one to pick the expert still does it but you just like reduce the time of expert by a lot lot so and and it will stay this way for really long so uh for the models to actually make good decisions they don't need to do correlation they need to do causation and as ya as Yan Leaken puts These are like text statistics not real world models.
You cannot just pattern match with past and use future to predict to it. And so we are like kind of not there yet. That's what I call as the AGI line. Once we are there, yeah, probably then models will just like make drugs. You will have VIP coded drugs. Someone would be VIP coding market. But yeah, not yet. So a final takeaway that I would call if if if there's one thing you are taking away from this talk, this is it. model infra ecosystem.
Everyone selling you tons of stuff at this conference is just commodity. Everyone has it. If you have it, everyone has it. Everyone can pay X number of dollars for a subscription. But what is mode and no one will come and sell it to you. You won't have to curate it on your own. Is the domain expertise. You need your data. You need other people's data. That is just not out there on the internet. And that that's what will form your mode.
So thank you for your time. I think you enjoyed the talk and yeah, let me know if you have any questions. We can meet outside. Thank you. [applause]
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.