Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

CodeTrading · @CodeTradingCafe
Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
14:184.0x the video's typical replay level
and I'm going to run it. So, it's going to take a while before it finishes. So, as you can see, you can follow up what's happening, the uh elapse time, the learning rate, the loss value. And once this is done, it's going to plot the
Said at 14:13
Most replayed moment #2
16:412.6x the video's typical replay level
this function, you might want to think how you can add additional technical indicators that matter for the trading that matter for the agent that can provide additional and valuable information for the agent to learn how to trade. One thing also that we could
Said at 16:34
Most replayed moment #3
17:532.4x the video's typical replay level
uh parameter. So, I just switched from 50,000 time steps of training down to 10,000 time steps. And this is the training equity again. It's working well. Now, we can test the model. The equity that we're going to get is
Said at 17:47
The graph counts replays. It does not show where viewers stopped watching.
Words
3,511
Runtime
19:52
Speaking pace
177wpm
Reading time
15min
177 words per minute, between the 160 25th percentile and the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Hi and welcome back. This strategy uses reinforcement learning to train a trading AI agent on historical data. I have chosen the URA dollar hourly time frame for today's example. The model showed an increasing equity on learning data and acceptable learning rate. In other words, we're trying to make an AI model that reads historical data, applies trading operations, and learns from winning and losing trades in order to come up with the best trading pattern pretty much like a human would learn trading. You can download the
89 words, the words spoken in the first 30 seconds at 177 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 218 |
| Average words per sentence | 16.1 |
| Longest sentence | 193 words |
| Questions asked | 2 |
| Sentences containing a number | 25 |
Most used terms
Filler phrases
96 in total: uh 63 · um 11 · actually 9 · like 6 · basically 3 · kind of 3 · I mean 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
Hi and welcome back. This strategy uses reinforcement learning to train a trading AI agent on historical data. I have chosen the URA dollar hourly time frame for today's example. The model showed an increasing equity on learning data and acceptable learning rate. In other words, we're trying to make an AI model that reads historical data, applies trading operations, and learns from winning and losing trades in order to come up with the best trading pattern pretty much like a human would learn trading.
You can download the Python code used in this video from the link in the description. Feel free to explore the files and use them in your own experiments. But first, let's summarize briefly how reinforcement learning works. Reinforcement learning relies on five main concepts. The agent, in this case, it's the one or the model making trading decisions, putting long and short positions, and setting stop-loss and take-profit distances.
The environment where the agent operates, the action or the actions that can be taken by the agent. Then depending on the result of an action, the agent receives a reward that can be either positive or negative with different values. One more concept is the policy which is the agent's strategy or mapping from different states of the environment to actions. In other words, it's the decision patterns learned so far by the agent considering the current environment state.
If you are familiar with reinforcement learning, this would be a model-free approach where the agent learns directly from experience or trial and error in the environment. Think about it as training a dog to complete certain tasks. Every time the dog does well, it will receive a reward. And if it doesn't, then either the reward is null or it can be also negative. Of course, AI agents don't have this intrinsic motivation for the reward.
So they can be coded in a way to maximize the reward value. In other words, if the agent receives more reward doing action A, then this action will be attributed the highest probability in the future. So now we will define our environment, the agent and the list of possible actions that the agent will use to trade in the created environment. And we will let the agent run trading on historical data for around 10 years for example and learning from the reward system the best possible ways of trading just like a human trader learning new strategies and experimenting with trial and error to improve his trading skills.
Okay. Now let's see how this is done in Python. It will be highly technical but you can still follow even if you don't have a purely technical or numerical background because the elements of the code are directly related to trading. I will show you the results and we will discuss how these can be used for our strategies. Since this project is relatively larger than the previous projects or studies published on this channel.
I'm not using a Jupyter notebook file this time. I'm using Python files because we're going to use few files like 1 2 3 four different Python files. So we're splitting the code into four different files in order to make it easier to read and easier to understand. And then we're going to launch it through Python. So we're not using a Jupyter notebook file in this case. However, when you download it, I will be sharing a folder with the data I will be using for this model.
So we can see that we have the Euro US dollar candlestick, the hourly uh data. It's between 2020 and 2023. And then I have another out of sample data or testing data where I've downloaded between 2023 and 2025. So this is the data on which we're going to test the model after we have trained it on the previous 3 years. I'm going to start with the um indicators.py file. So that's a Python file. I'm using pandas and pandas technical analysis.
It's just defining one single function here. Load and pre-process data. It takes a CSV path. So the path for the CSV file. We're going to read underscore CSV using the path. Clean the data somehow. Sort the index. add some technical indicators, the RSI14, moving average 2050, and the ATR, the moving average 20 slope as well. So that's just the difference between two consecutive rows. It's not really the regression slope on three or four rows.
So it's not really the slope, it's just a difference between uh two consecutive moving average values. And then we're dropping the uh empty values or the empty rows. We return a data frame with few additional indicators. And this is where if you intend to improve this uh code, this is where we can add some additional indicators and some custom indicators. Maybe you don't want to use the classic indicators. Maybe you want to add your own.
But I mean this is where we can do it. It's in this file. You can define whatever indicators you want uh as functions here and then use them to return the data frame with these indicators. Now we can define the trading environment and this is done using importing gym numpy because we're going to use it for numerical uh analysis or numerical uh operations and from gym we're also importing spaces the observation spaces and the action spaces.
So these are basically uh defining the actions that are allowed uh for for the uh the agent. And I'm defining a class named forex trading environment that inherits from gym.vironment environment class because this is how I want to create the environment by inheriting from the uh package gym. We define the uh constructor. Notice that we have a window size by default equal to 30. So by default the um uh the agent is going to read the last 30 candles before taking a decision.
And this of course includes all the related technical indicators that we've added in the data frame. Now, for the uh stop-loss options, I'm allowing the agent to take one of three decisions. Either 60 pips or 90 pips distance or 120 pips. Where do these come from? These are just random. On the Euro US dollar, I would put uh on the hourly time frame, I would put maybe 60 pips to 90 or 120 pips of stop-loss distance. Same thing for the take-profit options.
So we have 60, 90 and 120. You could add these, you could extend these, but the more options you give to um the agent, the more time it's going to uh take to train the agent and it's going to become computationally more expensive. So it requires more computation time. Now this shouldn't be a problem of course, but for the sake of this video, I kept it as simple as possible to show you how it's done. Now there are three options here for trades. either the action zero.
So it means that we're we're not going to trade. The agent is going to skip the candle without opening any trades. But in the opposite case, if action is positive or it's equal to one, then we have two different directions, either we short the market or the agent will long the market. So either zero or one taking into account of course stop-loss and take profits. We compute the shapes of the data frame to be used later on by the agent.
And this is where we initialize the uh current step which is equal to zero at first. We have the equity $10,000. Maximum slippage is zero for the moment. Then we have the uh positions list is empty at this point. And at the end I would like to um plot the equity curve. So we're going to log the equity curve values and the last trade information. The function within the class named get observation is simply to return the last window size.
Remember that we took 30 by default, 30 rows. So the agent is going to read the last 30 rows but also the number of features. So the shape is going to be window size. It's a numpy array by the way. So it's two-dimensional numpy array with the window size and also accessing all the features available on the data frame. It also takes care of the edge case at the beginning of the data frame when we don't have enough history.
So this is going to be used by the agent to read the data before each candle. And now we have to define the calculate reward. Remember that the agent is going to behave just like we are training a pet. Okay, we're training a dog. And we need to provide this reward function to know when we can reward in a positive value or in a negative value the agent. In this case, we're going to use the uh profit and loss for each of the trades.
So, if the agent opens a long position, for example, or any trade and the uh the result of profit and loss is positive, the reward is going to be positive. And it's also proportional to the uh profit and loss value. And that's why at the end of this function, the reward, which is the returned value, is the profit and loss times 10,000 because it has to be adjusted to the uh the number of pips uh from from the Euro US dollar price.
In the opposite case, if it's a loss, so if it's negative, in this case, we're going to return a reward that is negative. Now, since the data is processed in candles, you might have this extreme case where one candle touches the stop-loss and the takerit at the same time. We don't know since we're not using tick data, we don't know which of these uh levels were touched first, were triggered f first. So, we don't know if the trade was a loss or a profit.
In this case, just to be on the safe side, we're going to consider it a loss. So we don't cheat our way around. We're going to in case of a doubt, we're going to consider the trade it's been a loss and we're going to uh return a negative reward as well. And that's it basically. Now we still have to define one more function that's uh the step function. It's going to consider all of the previous uh functions. So we have direction, stop-loss, and a takerit and a reward and so on.
It's going to crunch it all together. and it's going to return the observation, the reward and other information. The reset function resets the uh the values of the back test. So, uh to the current step and the equity 10,000 and uh it empties the equity curve list and so on to uh to restart the test actually. Then we can render the equity curve using the render function. So that was our trading environment.py. Now to train the agent, we're going to use all of these functions and classes to make it happen.
So we're going to import from stable baseline 3 the PO. That's the um the model we're going to use to train the agent. And if again if you are coming from data science AI or reinforcement learning and you are a bit familiar with these, this is a model free approach. This means that the agent is going to learn directly from the environment. We don't have the whole environment map in front of us and we're asking the agent to learn from the map.
It's actually going to interact step by step with the candles with the historical data trial and error and it's going to adjust the parameters uh when it's progressing through the data. So first we're going to load and process the data. I have the euro US dollar hourly time frame 2020 2023. I'm creating actually an environment. So using the uh the class forex trading environment providing the data frame the window size of 30 the stop-loss options which you can change here the take-profit options then we have a dummy vector declaration that is required by stable baselines for parallelization so we're not going to uh go into the technical details but this is where you can also experiment to change the po now I'm not saying that you will be getting better results it's just that for learning purposes and for educational purposes.
It's good also to train the model using different models or uh the agent actually train the agent using different models. The PO is actually one of the best suited models for uh for this case for financial trading uh simply because this is a very noisy and continuous data. So we assume that the data is flowing continuously. you the agent is going to receive data live from the market and it will build its skills trading live on the market checking whenever we have a loss whenever we have a winning trade and so on so it's kind of acquiring this experience and in this case PO works well then we're going to train the model so using modelarn the total total time steps actually is 50,000 you might want to adjust this one as well if you increase it it's going to uh train the model more in a better way but it's requires also it requires also uh additional computational power then we can save the model as model euras dollar and that's the zip file you can see here I've already run and trained the model and I've saved it here for later to be used and uh at the end we print model saved successfully then we can evaluate or test the model so we're going to um reset the environment uh the equity curve is empty again and then we're going to try to predict use the model to predict values using the observation space.
We don't have any stochastic uh operations included. So that's deterministic and for each um step we're going to apply the action buying or selling and so on. We're going to record the observation, the reward and if it's done or not. So if it's finished or not and some additional information and at the end I'm going to append the equity uh with the current equity actually. So that's going to build our equity curve. Uh so we're going to append all the equity values or the balance values in the equity curve list to be able to plot it using this part of uh of the file.
To be able to run this, I will be opening a new terminal. I'm going to um source my environment because I've created a virtual environment specifically for this application. Virtual environment scripts activate. I'm going to run python train uh agent.py. pi and I'm going to run it. So, it's going to take a while before it finishes. So, as you can see, you can follow up what's happening, the uh elapse time, the learning rate, the loss value.
And once this is done, it's going to plot the equity curve on the training data. And as you can see, it's training well. So it's managing to increase uh the equity which means that basically the agent knows uh that it needs to follow the positive rewards action or course of actions. So it's building a kind of policy that will allow the equity to increase uh over time with it with the different trades. Now the correct way to evaluate the agent is actually to test it on new and unseen data.
This is the training data and it's working well. We can see that the equity is positive. the agent is heading in the correct direction. Let's close this one and head to the test agent.py file. And here I'm loading a new unseen data. So that's always the uras dollar candlesticks 1 hour uh time frame. But this time it's 2023 up to 2025. I'm not going to train the model here. It's not going to fit or apply any fitting. It's just going to try and trade.
But I'm still using the same options for stop-loss and take profits. The same window size as what we have used for the training part. And I'm going to run this uh same model. So we're loading the model that we have saved from the uh training phase or the learning phase. And we're going to try and test the agent. So python test agent.py. It also takes a bit of time before showing the equity curve. And this is what we're getting.
So it's not what we've expected. The equity was showing a more positive trend using the training steps or the training data. But this is on new and unseen data. It's not very bad either because to be honest we didn't provide much for the um uh for the agent. So these are barely few technical indicators. These are classic technical indicators. the two moving averages, the RSI, the ATR, and what we're calling a slope, which is not really a slope.
We didn't apply aggression. This is just to show you an example as simple as possible actually just for the sake of this video. But in this function, you might want to think how you can add additional technical indicators that matter for the trading that matter for the agent that can provide additional and valuable information for the agent to learn how to trade. One thing also that we could be changing in the trading environment these options.
So I've also provided very limited options for the stop- loss and take profit values. So it's either 6090 or 120 pips for both. Maybe this is not enough for trading the Euro US dollar. maybe we should be providing more values maybe with five pips of increments uh covering from 30 pips up to I would say 100 pips for example for both stop-loss and take profit options so that's one thing to be uh to be improved also when we were training the agent we used 50,000 time steps maybe the agent is overfitting we could try to um fit it with 10,000 steps for example so now if I'm interrupting ing and retraining with 10,000 steps.
I'm going to show you, let me save it first. So, I'm going to show you the uh effect of this small uh parameter. So, I just switched from 50,000 time steps of training down to 10,000 time steps. And this is the training equity again. It's working well. Now, we can test the model. The equity that we're going to get is slightly different. So as you can see it's kind of positive at first, it's not as positive later on and so on.
So sometimes you might get the set of parameters that will allow you to um train the model but then test it also on unseen data and you can still have this positive equity trend uh even when you are testing on unseen data. So as you may have noticed there's a lot of options that we can still use to improve the trading agent uh potential. I still find it enjoyable to be honest. Uh adding some more stop-loss options, take profit options and trading indicators, maybe the classic or custom indicators and so on.
Changing the total time steps uh in order to avoid overfitting the model on the training data and so on. So, it's a lot of fun and seeing the results straight in front of us is also powerful because you can see the effects immediately in front of you. And that will be it for this video. I hope you guys liked it and found the information helpful. I've been receiving lots of requests recently. Why don't we use the reinforcement learning in trading?
As you may have noticed, it's very powerful. It has a lot of potential, but it's not as simple to fine-tune as well because you have a lot of noise in the trading data. So this is the nature of the market and sometimes the model is finding it difficult to pick up the signal the true signal of the trend from all the noise that you can find in the data. But anyway I hope you guys liked it. I hope you found this information helpful.
If so please leave a like, leave a comment, drop your ideas in the comments section. This was requested by one of the uh viewers through the comment section. So don't hesitate to drop us some of your ideas. Thank you so much for watching. Thank you so much for staying that long. Until our next one, trade safe and see you next time.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.