Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Alejandro AO · @alejandro_ao
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in Alejandro AO's most watched videos.
Most replayed moment at 28:46
6.8x that video's typical replay level
the binary going to back go back to my um Cloud config file and instead of just saying UV I'm going to add the entire binary path right here so now I can close this and I can restart
Said at 28:39
Most replayed moment at 5:31
4.0x that video's typical replay level
In my case, as I told you before, it is front-end design.
Said at 5:36
The graph counts replays. It does not show where viewers stopped watching.
Words
5,929
Runtime
35:39
Speaking pace
166wpm
Reading time
25min
166 words per minute, between the 160 25th percentile and the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Good morning, everyone. How's it going today? Welcome back to the channel. In today's video, we're going to be covering how to prompt your AI coding assistants, such as Trey, Cursor, Copilot, Windsurf, whatever you're using. And we're going to be seeing how to prompt them to actually create useful and secure applications that are robust and that do not acquire infinite technical debt that will make you want to pull your eyes out. So what we're going to be covering today is
83 words, the words spoken in the first 30 seconds at 166 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 322 |
| Average words per sentence | 18.4 |
| Longest sentence | 89 words |
| Questions asked | 20 |
| Sentences containing a number | 11 |
Most used terms
Filler phrases
125 in total: actually 88 · like 25 · I mean 6 · kind of 3 · right? 3.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, published by the channel, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
Good morning, everyone. How's it going today? Welcome back to the channel. In today's video, we're going to be covering how to prompt your AI coding assistants, such as Trey, Cursor, Copilot, Windsurf, whatever you're using. And we're going to be seeing how to prompt them to actually create useful and secure applications that are robust and that do not acquire infinite technical debt that will make you want to pull your eyes out.
So what we're going to be covering today is how these AI coding tools work, so how Cursor works, what is actually happening behind the scenes, and how you can think about this tool so that you can better use it. We're going to talk about what makes a good prompt. I'm going to show you actually some examples of pretty long prompts that are very, very useful to start actually a thread or a conversation with your assistant.
We're going to be learning how to see conversations as threads of tokens, essentially, so that we use them in the way that they're actually working behind the scenes. And we're going to be covering a very quick demo of how to actually create one of these prompts and execute it. So without any further ado, let's actually get right into the video. All right, so now let's talk a little bit about how language models work.
And this is especially important because this will give you an intuition of how to use these AI tools that we're working with and that you're most certainly going to be working with if you are not already. We have already talked about this a little bit, and we talked about how language models work when creating AI applications, but this intuition is also very useful to have it when you are using these AI applications, okay?
So in this example, we have just a very quick diagram of how language models work, and this is a way to think about it, the best way to think about it, I would say, when using these tools. You have to think about a language model or these AI coding tools as essentially just a text-to-text machine, all right? So essentially what's going to happen is that you're going to send some text to it that is going to be your prompt.
This text can include just a question, it can include a task, a very detailed instruction of a task, or it can contain just a short prompt and then a very long piece of code that it has to complete, etc. And then as the completion, you will have the most likely continuation to whatever you sent right here, okay? Now, how does this actually work? To give you a little bit of an idea of what is actually going on behind the scenes, the language model is actually not outputting a single token every single time that it receives an input.
Actually, the output looks more similar to a distribution. So the final layer of your language model will generate a distribution of the most likely tokens that come after whatever input you give to it. So for example, let's suppose that you give to your language model the input, why did the chicken cross the, okay? It's a cut-off sentence, it's not even a question. Now, the language model is going to return a distribution of the most likely continuation to whatever you sent right here.
So it can say that road maybe has a 20% probability of being the most likely continuation right here, then river, 10%, house, 5%, banana, planet, etc. And you'll have all the way to all the words in its dictionary, and all of the probabilities right here will total a 100% probability, okay? So that is actually what is going on right here. You're generating a distribution of the most likely tokens, and this one has a 20% probability of being chosen as the next continuation right here.
Now, there are different ways to organize this distribution and to make it more or less spiked. Usually, it is the most likely continuation has a significantly higher probability of being chosen than the rest, but that's not always the case. And there are ways to tweak this, but usually you have to worry about it. It's already done by default. And also something to keep in mind about this is that this is mostly true or more visible with older models.
Newer models are actually fine-tuned to follow a conversational system, conversational manner, in which if you send something like a cut-off sentence, instead of actually completing the sentence, they will continue the conversation saying, hey, I guess your message got cut off. Can you complete whatever you were sending? But in the end, what they're doing is just completing the sequence of tokens with the most likely continuation.
And that's true not only when you're talking to JGPT, to Anthropics Cloud, to whichever assistant. It's true also for coding assistants that use these language models behind the scenes. Everything that is happening in the language model is going like this. And the input can be a very long piece of code, and the continuation will be the most likely continuation according to whatever you requested from that piece of code.
So there you go. That is a way to think about language models and to think about code. And they will be generating the most likely continuation. All right. So let's talk about how to create a good prompt. And it's not as straightforward as it sounds. It's actually more of an art than a science. But right here, I'm going to show you some tips, tricks, some advice to, in general, help you create better prompts, not only for your AI coding assistants like Cursor, Trey, Copilot, but also for pretty much any other AI assistant, such as JGPT, Cloud, etc.
And the first thing that you have to consider is that you have to shift the way you think about prompting or about querying to solve a problem. We have, for a very long time, been used to actually using Google to figure out solutions to our problems. So, for example, in Google, say you wanted to enter some data into your database, you would search for something like how to enter data in SQL. And this is already quite a good query for Google, but this is probably not the best for an AI.
So an AI would probably take something better, like write the SQL query to enter this data to my SQL database. And then you can send to it something like my data is the username, the username is whatever, and then the email is whatever. And then the table name is this right here. As you can see, this is already becoming more of an actual instruction that you would give to someone who's actually building this, a human, rather than just learning how to do it and then trying to figure out how to put everything together.
Because if you send something like this to your language model, be it on Cursor or in Copilot, whatever, it is just going to respond to you with a very simple query, and then you're going to have to update it by hand, and you will have lost time. You will have wasted time by waiting for the system to generate the response and then tweaking the response. Whereas if you had sent something very precise from the start, you probably would have something much closer to what you actually needed.
So just a very quick example, but a good way to think about this in general is to think about it as though you were talking to a human that is very, very good at overly making assumptions about everything that you say. But it's also very good at executing things. So let's suppose that you go to, an example that I really like is let's suppose that you go to a travel agency. And in the travel agency, you arrive and you want to travel to China.
So you tell your travel agent, organize a trip to China for me. Now, a human will most likely come right here and be like, okay, yeah, sure. Let's organize a trip to China for you. So where is your origin? Where are you going to live from? So you're going to live from San Francisco, maybe. What is your budget? When do you want to go and when do you want to come back? What cities would you like to visit, et cetera? So these are all of the questions that the travel agent is going to ask back at you.
That is not how a language model is going to work. A language model is actually going to make assumptions about all of this, because they will be the most likely things to come after it starts writing the recommendation for a China trip for you. So it will give you the most likely continuation for the origin. It will give you the most likely continuation for the budget, the most likely continuation for dates, et cetera.
So what you want to do, I mean, depending on how fine tuned it is. But in general, a good thing to assume is that it will make assumptions about everything. So you have to be very clear about everything that you want it to do. So actually, if this was a language model instead of an actual history, human, you would send all of this from the start. OK, that's a good way to think about it. So all of the instruction that is very precise has to go right there from the start, everything from you.
OK. And now I understand that it is not very straightforward to do. It's not something that's very straightforward to do when you're trying to create something, especially in software. Sometimes you don't really have a very clear idea from the start of exactly where everything is going. So if you're talking to an assistant, maybe you're trying to figure things out at the same time as you're talking to your AI. And we're going to get there in just a moment, how to actually figure all of these things out and how to start a new thread with whatever you found out.
But something to keep in mind about this is that a good rule of thumb when creating this prompt is to add, is to consider this three pillars of a good prompt. You can consider that you have to have a very clear instruction, as we mentioned, a very clear context. So where you're coming from, what do you want, what you want to do, what you travel for, etc. And very clear examples also of what you want. This can be examples of exactly what you want the language model to generate, the format that you want the output to be, etc.
I'm going to show you some examples in just a moment, but just consider these three things. You have to have instructions, very clear instructions, very clear context and very clear examples. OK, and you may see that this is already becoming quite long. And that is true. A very good prompt is a very long prompt. So actually, let me show you an example of a prompt that I built just a few days ago for precisely for YouTube.
So let me show you that. All right. So as you may know, I do YouTube videos. So I actually needed the other day a tool that would allow me to concatenate all of the sequences of short clips that I record to put them all together into a single video for YouTube. And usually I do this manually on CapCut, go there real quick. But what happens is that I'm trying to make more videos, so I'm trying to make this a little bit more automatic.
And since I don't do a lot of post-processing, I felt that I could just write a very quick script to concatenate the whole thing and generate a few things that I need every single time that I create a video. And I figured that I could make a script that would look something like this. So actually, the script has a few methods, as you can see, it has a concatenate video, generate timestamps, generate the transcript, generate the description, and then a few helper functions and then generate the SEO keywords.
And it's quite a long, quite a long class right here. It is 432 lines right here. And the entry point is over 75 lines. So in total, about 500 lines of code. And I ended up coding the whole thing in about an hour. And the reason for that is that I used AI coding tools and I was very precise in exactly what I wanted to get. Okay. And actually the instructions that I used are right here. So this right here is the prompt that I sent to my language model.
So that it could understand exactly what I wanted to get. And I sent this one to cursor, if I remember correctly, or maybe at the time I was using Trey. But in the end, it is a language model that is doing all the work behind the scenes and it's working through my AI coding assistant. So the instructions look something like this. So there is a very clear overview of what my tool, the entire thing that I am trying to build actually is.
So this tool concatenates multiple MP3 files. So this tool concatenates multiple MP4 videos from a directory. It generates timestamps, creates transcripts, etc. I specify the configuration and the output of the organization, how the output is going to be organized, sorry, and the actual operations. So I have one, two, five, five operations. And every single operation, as you will see, is actually very well defined. So for each operation, I described a very clear objective, the requirements that I needed it to have, the steps that I needed it to take exactly.
And for example, for the second step right here, the objective was to create a JSON file with the chapter information of each video segment in my video. And I wanted it to be exported to JSON, but I wanted it to be in a very specific format so that I would be able to parse it afterwards. And so in order to do that, I just gave it a very clear example of how it had to parse the whole thing. And that's how it goes. It's very straightforward.
So this part right here is what I mentioned before when we came to examples. It's not only examples about what technology you're using or what the final code should look like. It's also examples about the schemas and the structure that you expect to get. So this is just one example right here, just an extra information right here, more information right here. And here's another example, as you can see. So this part right here is create a VTT transcript using OpenAI's Whisper API.
And the steps are very clearly defined. And something that happens sometimes is that the language models actually use previous versions of the SDKs that you're trying to use or the libraries that you're trying to use. So a good rule of thumb to do when you're prompting your AI assistant to create something with a very specific version of a library or even if it's just the latest version of a library or a framework, a good thing to do is to send your language model an example of exactly how the code looks like of what you're trying to build.
And what I did in order to get here was very straightforward. I just went to OpenAI Whisper API, I looked for the documentation and I copied the example right here with the latest implementation of the API, of the SDK. And there you go. Then you have more details right here. So as you can see, it's kind of programming in natural English language. And that's actually what we are doing. We are programming the language model, telling the language model what to do.
And that is why it is very important to keep all of these things in mind. So you have, as you see, we have a very clear instruction. You have a very clear context and you have very clear examples of exactly everything that you need. The examples are particularly important because your language models are not trained on everything, especially not new material. So if there was new material that came out in the past few weeks or in the past few months, the latest language models probably don't even have it because they are pre training material and that before that new material came out.
So a good thing to do is to send them an example of the new SDKs or the new documentation that you're trying to use so that they can actually absorb that on the spot. Language models are actually very good at absorbing information from your context, even if it was not in their training material. OK, so just keep that in mind. So, for example, if you're creating a script with a LangChain, LamaIndex, whatever, a good thing to do is to go to documentation of whichever library you're trying to use, copy an example of the implementation that you're trying to do, paste it into the context of your language model instruction and then ask for it whatever you want it to actually do.
All right. So there you go. So now let's actually take a look at how I actually made this prompt because it's a pretty long prompt and it can take quite a bit of time to actually get here. So let's take a look at that. All right, so we are going to see now how to create a prompt that looks something like this, which is a very specific and very clear prompt. Remember that we mentioned before that the larger the prompt, usually the better because it will have more context to actually perform the correct calculations to give you a correct completion.
So a very long prompt is actually exactly what you want to create. But how do you get there? As you can see, the prompt that I have right here is actually pretty long and it takes quite a bit of time to actually get to this part right here. In itself, it's actually kind of programming it in English. So how do I get to actually creating a prompt that is this structured? As you can see, it is quite well structured. And the reason for it is that it was actually a language model which wrote the final draft for me.
It was not myself. And that's what I want to show you right now, how to use a language model to actually get to this place right here before you can actually start building with your AI assistant. So when you start the process, I would usually start with an assistant that is already within my code base. And to do this, I go to cursor, for example, I come right here. And in this example, I already have a code base that is open.
But most likely then, I mean, you will most likely be in either of two scenarios. So you're starting a project from scratch or you're starting a project which is actually a feature of an existing code base. if your situation is situation number one, you're starting a project from scratch, I would actually encourage you to start the conversation, not with your AI editor, but instead with the chat GPT voice mode or your AI assistant of your choice, because that way you're able to actually just talk to it and describe your project and it can reply to you in voice and you can continue the conversation and it will be just like you're brainstorming your ideas about the project that you're trying to build.
And something that is important is that whenever you're doing this, always tell your AI assistant that you're in the planning stage, okay? So you're brainstorming, you're figuring out how to implement different features so that it can also give you advice or suggestions about how to implement whatever you're trying to implement. And then by the end, after you have been talking to your assistant for a good 20 minutes or whatever, after that, you ask your assistant, okay, so now we have a very clear idea, you and me, of everything that my feature or my application should have, now help me write a prompt that will be used in a new AI assistant to create this application or this feature that we have been talking about.
Describe it in a very precise and detailed way so that the AI knows exactly how to implement it. Describe the context and also give examples of whatever we have been talking about. So that's what I would expect you to do. So for example, let me show you. If you go right here, let's say that you are already working in a code base that already exists. If it is a code base that already exists, instead of using your preferred AI assistant such as chatGPT voice mode, I would actually encourage you to go straight into your ID coding assistant because it will have automatically all the context of the existing code base.
So right here, I'm working with an MCP client. So let's suppose that I want to implement a new feature. So I will ask it to implement this new feature. I am trying to implement a new feature to this MCP client that will be able to not only generate the completion from the language model, but that will also stream the completion as it is being generated. For context, this application is an MCP client, which means that it works with the model context protocol by Anthropic.
And essentially what it means is that it is able to connect to multiple tools that are hosted in this thing called the MCP server. As you can see, there are some methods that already connect to this MCP server and that get the tools. All I want you to do right now is to actually just take the tools that were generated, execute them if there are any in the language model response. And then after that, when the language model actually finishes using the tool, you will have to stream the generation token by token with a generator that I will expose via my API.
And there you go. So as you can see, I mean, I described pretty much everything that I wanted from my project and actually getting a dictation tool, something like Super Whisper, Willow or Open Super Whisper for that matter. It's also quite useful because you can talk about what you want and you can be a little bit more detailed and go around whatever you're trying to build. And as you can see, it took me not too long to actually write a super long paragraph.
And I can actually continue this by saying, help me plan this feature and make sure that everything that I want to implement is actually correct and cogent with whatever there is in my current implementation. I want to make sure that everything works correctly. Please ask me questions if I forgot to mention something, if you feel like something needs a little bit more detail and let me know if there are some other libraries that I should implement in order to actually implement the feature that I am trying to build.
And there you go. So this is my entire prompt. I'm gonna send it. And now Cursor is going to start actually thinking about the whole thing. And it's going to consider the process that I already have right here. And we can go back and forth. And right here, this is a planning stage. So we're going to be talking with Cursor. We're going to be identifying what I need to actually do, what actually was not a good idea, probably that I mentioned before.
And then by the end, after a very long sequence of back and forth, I can just ask it. Perfect. Now write a very detailed prompt that I will use with a language model. And this prompt will be used to create and implement this new feature that we have been talking about. Make sure to make it as detailed as possible and include examples whenever needed so that the language model knows exactly how to implement the feature that we have been discussing.
Then there we go. And now the language model is going to be able to actually generate the prompt for me. I can actually ask it to generate the prompt inside a code block in Markdown. And there we go. Now we're going to get the actual prompt right there. And this is the actual prompt that we're going to be able to use within my application, whichever it can be. It can be Cursor, it can be Tray, it can be Windsurf, whatever you want.
And this is exactly how I built the prompt that you saw right before. So as you can see, it's a pretty long prompt. It's a pretty detailed prompt. I did not go over the whole back and forth right here, but here we already have a new starting point for my next part right here. But what I want to tell you right now is that you should copy this prompt right here and actually start a new conversation. Because what some people do is that after talking about the feature for a while, they start talking about, let's now implement it.
And sometimes that actually works pretty well. But I'm here to tell you that there's actually a better way, which is just copy this extracted material, this extracted prompt from the entire conversation and start a new thread with it, okay? So what I would do is right here, I would come right here. I would copy this prompt right here. I will start a new conversation right here with Cursor. And I will start it with this prompt right here.
And then just maybe add a little bit of extra content before that or afterwards. And this is going to give me much better results. And as you can see, it is already writing the code and implementing it. And now I would have a much, much better implementation of what I wanted to build. So as you can see right here, I have my new methods that were created. And this is how I would actually approach the prompt engineering part right here.
And there you go. So it finished generating the whole code. And something that is very important right here is to actually read the entire thing that was generated to make sure that it actually makes sense. This is especially important, especially if this is not just a very quick weekend project that you're trying to vibe code your way through. But if this is a real project that is actually going to production, you want to really understand the code that is being generated because otherwise you're going to just accumulate a lot of technical debt.
Technical debt means that you're going to at some point have to come back here and understand the code, especially if there's an error. So if you understand the code from the start and are able to actually update it to whatever you actually need, because as we saw before, the language model will not always generate the best code. So if you actually read it, understand it, and actually tweak it to make sure that it is exactly what you need, you're going to be actually using the language model to save you time and not to build technical debt, which can be a hell because language models can be very good at building up a lot of technical debt for you.
So just keep that in mind. Use the language models and use the code that is generated by them as though it was your code. Keep it to the same standards as you usually would with your own code and you will have everything you need to actually go faster into your development. But now I wanted to mention real quick why actually I did what I did with which was go through the conversation and then restart the whole thing again in a new thread.
All right, so just very quickly, let's talk about why it is important to do what I showed you just now, which is starting a conversation within a new thread with that new, very detailed prompt that you have just created. And I mean, you do not always have to do it. Actually, if you have started your conversation with your assistant and a few messages in, say six messages in, you already have a very clear idea of what you want.
You can just go into implementation right here. You can just ask it, switch to agent mode and ask it to implement it, implement the feature. And that's more than okay, all right? However, sometimes it happens that when you're exploring the idea, you're exploring the feature, you're discussing with your assistant, everything, you can, I mean, this can actually take more than just this. I mean, you can take like say 20, 30 messages or even more, and that just creates a very, very long thread, a very long conversation.
And something to keep in mind is that every time that you send a message to. to CGPT, to Cursor, to Copilot, to any generative AI that you have right here, what you're actually doing is you're sending the entire history of your conversation on each message, right? Because the language model itself does not have a memory of the conversation. You have to feed it the entire conversation every single time to make sure that it knows what you were talking about.
So actually, what happens is that every message is sent every time. And this starts even from the first message. When you start a conversation with your assistant, that is not actually the first message that is sent to the language model. The first message that appears in the thread is the system prompt, okay? And I'm sure that you have probably already heard about the system prompt, what it is. It is essentially the first message.
It's usually a system message that describes the personality and the objectives of the assistant or the AI that you're talking to. And for example, in the case of CGPT, it's something like you're a helpful assistant, et cetera. They do not really release their system prompt. In the case of Anthropic, let me show you, Anthropic system prompts, they have actually open-sourced the system prompt. So if you go right here, you can see the system prompt from Claude.
And this one is actually the first message that appears in every conversation that you have with Claude. And that is actually what I want you to think about right here. When you're talking to Cursor, you're not only sending your own message to the language model, you're sending the message and then that is being appended to the system prompt that the application created. And then after that, once the AI responded, it responded with its own AI message, then you respond with your own message and et cetera, and you continue the conversation.
And as you can see, this is kind of a thread of tokens. So this is the first token and the first set of tokens. Then you send another set of tokens and the AI continues the completion with another set of tokens. And then you continue the thread with another thread of tokens, et cetera, all the way until your entire conversation. And remember that all the history of the previous tokens in the entire thread are used as context to generate the next response.
So all the thread is actually being sent right here as input every single time. And that's why I said that it's sometimes a good idea to actually restart the thread, especially if you took a long time to actually figure out what you actually wanted in the feature. So let's suppose that you have been talking to the AI for like 40 messages about what you want to build. And in the meantime, you went through tangents and you considered features that in the end you will not be including.
You considered other libraries that you decided not to use, then chose a different design pattern with the AI that you wanted to implement, et cetera. All of that will be in the history. And if after all of that, right here, you arrive and you're like, okay, so now implement it, it will have all the noise of the previous conversation. So that's why it's a good idea to just ask it to give you a very clear prompt like the one we had before with a very clear explanation of exactly everything.
So a clear, long prompt that the AI is going to generate for you. So I'm going to make this a little bit longer like this one. And now what you can do is take this one and go up here and start a new thread from that new system prompt and start a new thread like this. So that way the thread is really focalized on your specific feature that you want to implement. All right. So that's why I recommended that you restart the thread, especially if you go very long right here.
So there you go. That is how you build prompts that can actually help you build applications faster. Just remember to think about your prompts as though you would think about an actual program. You're actually describing everything that you want your application to do in your prompt. Okay. And the longer the prompt, the better. Remember that you're not talking to Google. You're talking to a language model that will take everything within your prompt to give you the most likely completion, which is going to be the code that it will generate.
Okay. So just keep that in mind. I hope that this video has been useful. There are, of course, many different ways to prompt your language model. And sometimes a long, super long prompt is not the way to go. Sometimes you will have to go for shorter prompts that are very specific to whatever you're working with at the moment. But I hope that this has been quite useful. This is an approach that I see very rarely taken and that I think it's very useful because it will increase the accuracy of your language model in the generation of your feature.
So thank you very much for watching and I will see you next time.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.