Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
![I Let GPT 5.6 Sol Agent Trade Gold for 50 Days [Full Backtest]: video thumbnail](https://i.ytimg.com/vi_webp/lrLdK6t3M5w/maxresdefault.webp)
CodeTrading · @CodeTradingCafe
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in CodeTrading's most watched videos.
Most replayed moment at 14:18
4.0x that video's typical replay level
and I'm going to run it. So, it's going to take a while before it finishes. So, as you can see, you can follow up what's happening, the uh elapse time, the learning rate, the loss value. And once this is done, it's going to plot the
Said at 14:13
Most replayed moment at 14:06
5.2x that video's typical replay level
predictions in the bearish category as well. And we have category zero as well, so it's populated with a precision of 68%. So, this is how we can play with the model parameters, prediction probability, the hyper parameters, and
Said at 13:59
The graph counts replays. It does not show where viewers stopped watching.
Words
2,646
Runtime
15:40
Speaking pace
169wpm
Reading time
11min
169 words per minute, between the 160 25th percentile and the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
I let GPT 5.6 Soul trade gold for 50 days, but I blindfolded the model so it never actually knew it was trading gold. It made 300 decisions using a paper account, and in this video I'll show you how I did this, what are the pitfalls, and of course, we will go through the results. A few days ago, OpenAI released GPT 5.6 Soul, its new flagship model. So, I gave it a trading account price level index with some indicators like the RSI,
85 words, the words spoken in the first 30 seconds at 169 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 204 |
| Average words per sentence | 13.0 |
| Longest sentence | 59 words |
| Questions asked | 3 |
| Sentences containing a number | 36 |
Most used terms
Filler phrases
30 in total: uh 13 · actually 5 · like 5 · I mean 2 · kind of 2 · um 2 · you know 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
I let GPT 5.6 Soul trade gold for 50 days, but I blindfolded the model so it never actually knew it was trading gold. It made 300 decisions using a paper account, and in this video I'll show you how I did this, what are the pitfalls, and of course, we will go through the results. A few days ago, OpenAI released GPT 5.6 Soul, its new flagship model. So, I gave it a trading account price level index with some indicators like the RSI, ATR, and the exponential moving average.
Then all was fed into a trading agent that reads the market data and makes its own decisions. Then it chooses whether to buy, sell, close, or wait. It also chooses the stop loss and the profit target. And after every decision, it actually explains why it made that choice. I also included a test that most AI trading videos ignore. We will come back to that problem soon. The full Python code is available. You can download it for free from the link in the description of this video.
I will also try to add it in a pinned comment. It also includes a mock mode so you can run the project without using or paying for an API. Okay, so here's the experimental setup. I used 60 days of gold data from CCXT. The data uses the 1-hour time frame, the 1-hour candles, and also daily candles for more context. The first 10 days are only used for context or to give the agent some market history. The agent does not trade during those 10 days.
It then trades the next 50 days. It makes one decision every 4 hours. That gives us 300 decisions in total. The starting balance is 100,000 paper dollars, so I'm using a paper or a demo account. I also included trading costs, fees, and so on to make this test as real as possible. The test uses five basis points for commission on each side of a trade plus two basis points for slippage. The agent can go long or short. It must also choose a stop loss or a take profit target, or decide when to close a trade.
All of it is running in a back testing style approach using the 1-hour time frame. Now, here's everything the model receives before each decision. First, it sees the last 24 completed 1-hour candles. It also sees the 10 completed daily candles. These daily candles show the market from a wider view. This helps the agent understand both the short-term and the longer-term trend. The model also receives several indicators.
These include the RSI, the distance from different EMAs, for example, and the ATR. All of these values are shown just as percentages, not raw values. Remember that point because it becomes important later. It's mainly part of the blindfolding approach to blindfold the agent, and I will explain more in details why is that important in a while. The model also sees its current account information. It knows its balance, its open positions, and its recent performance.
Finally, it sees its own recent decisions. In other words, the agent reads its own trading journal before making the next choice, just like any normal trader or a human trader. After reading all of these or this information, the model must make exactly one function call. That function is called place {underscore} order. It must choose one action, buy, sell, close, or hold. It must also choose the position size, the stop loss distance, and the profit target or the take profit distance.
Finally, it gives one short reason for its decision. That reason is saved directly in the trading log, and we can access the trading Here, for example, I can go into agent runs. I go into this run, and I can, let's say, the decision terminal log, for example. You can see that for this case here, we have hourly rebound is constructive above 20 50 EMAs, for example, but price is stalling and so on. So, you have those messages that is showing for every decision why is the agent acting, why is it holding, why is it entering the market as a short or buying or closing the trade and so on.
So, everything is explained step by step what is the trader or the agent trader doing. It's also logged here. I also included five protections against look ahead or future data leak. First, the agent only sees candles that have already closed. The data frame that is sent to the AI agent is physically cut at the current decision time. Anything after that point is hidden to avoid looking into the future before making a decision.
Second, the program uses only the visible daily data. If the current daily candle is still forming, it's not totally closed, so it is removed because it didn't close yet. The agent only sees completed daily candles. Third, an order is not filled at the price the agent just saw. It is filled at the opening price of the next 1-hour candle. Then, every indicator is calculated using only the data available at that moment, meaning using only closed candles.
The same rule applies to ATR, which is used to calculate the stop loss distance. Those are the first four protections. The fifth protection is the main reason I created this experiment. Here's the problem with many ChatGPT or large language models trading experiments. Large language models are trained using information from the internet and the internet contains a huge amount of historical market data. It contains charts, articles, trading discussions, and complete price histories.
So, for example, if we were to backtest an AI agent using candles from 2024, the AI appears to predict what happens next, but this is probably due to the fact that it recognizes a price pattern that it had already seen during training. There is usually no way to know, but this means that some impressive AI trading results may come from memory and not really from um real trading edge. To reduce this problem, I used three controls.
The first control is time. GPT 5.6 Soul has a knowledge cutoff of February 16th, 2026. Our trading test period begins on May 21st, 2026. That is 94 days after the cutoff of the training data for GPT 5.6. So, the model could not have seen this exact market period before. The second control is much stronger. I completely blindfolded the model. If we look at the actual prompt, it never says that the asset is gold. It never shows the ticker XAU.
The real prices are also removed. Instead, the price becomes an index that begins near 100. The real dates are removed, too. The model sees time labels such as T plus zero days and eight hours and so on. Volume is not shown as a real number, neither. It is shown as a multiple of its own recent average. So, almost every value is converted into a percentage or another scale-free measurement. In brief, the model does not know the asset, the date, or the real price.
It cannot remember a market that it cannot identify. The third control is the benchmark. I compare the model with buy and hold. I also created 20 random trading agents. These random agents use the exact same backtesting engine. They pay the same costs. They use the same stop-loss and take-profit system, but their decisions are based on a random coin flips. So, these are totally random decisions. If the AI cannot clearly beat this random group of agents, then it's simply as good as a random agent.
I will not go through all the Python scripts in detail, but I just wanted to show you some of the files. So, there is the config file here, config.py. This is where you can change most of the configuration or the parameters. Then, the most interesting part is the agent actually here. You can read the prompt, which is right here at this level. So, this is what I wanted to show you. This is what we pass to the agent. You are trading asset X, so it's not known.
A liquid market, its identity, uh calendar dates, and price scale are deliberately hidden. Prices are an index. History starts near 100, and times are offsets from the start, etc., etc. Do not try to guess which asset or which time period this is. Reason only from the provided market snapshot. Then, we have the rest of the prompt. You make decisions on 1-hour candles, and also receive daily candles, etc., etc. Only completed candles are ever shown to you.
Uh account rules, we have one position at a time, the close order execute at next 1-hour open, etc. So, each decision cycle, uh review the complete market snapshot, reason briefly under 120 words. So, let's keep the uh reasoning uh short as short as possible in order not to consume tokens either. Remember, this is a very uh fast uh repetitive uh work, so it can blow your uh budget very quickly using the API. Call place order the function exactly once, buy, sell, close, or hold, and so on.
Uh be selective, holding is often the right call, but we still need to enter the market in the direction of the momentum at the right entry moment, etc., etc. So, this is where you can change things to make the agent act a bit more aggressively, but with this was not the point of this experiment. I just wanted to show you how it works. Then we have the main script here, main.py. This is where you have some of the comments showing you how to run the experiment.
We can use synthetic data, which is generated locally, and the mock agent, so you don't have to use an API or using anything that you would pay for, but this is of course all mock and synthetic run just to make sure that the experiment is running and how the export or the output results would look like. Then you can also, if you have your CSV file of data, you can use this using this part. And if you want a complete normal run using the API key mainly of OpenAI, you can run python main.py.
Then asset is gold, so this is where it's going to get the gold from CCXT. You don't have to put CSV, etc. So, if you don't put it, it's going to download the data from from the CCXT repo. And then it's going to run a real agent using OpenAI using OpenAI models. Now, I think in the config I put by default the model is GPT-5.6 small, as you can see here, but you could use something else actually. If you want to have a more economic or lower cost alternatives, you can use GPT-5.5 mini, for example.
And the results are saved here in the runs. So, we can see that we have decision log, we have the interpretation of the market by I mean the agent, how it explained the market, and what kind of decision it took. So, for example, this is hold, the equity is a hundred thousand dollars. Let's see here, hourly rebound is constructive about twenty fifty MAs, but price is stalling, for example. So, this is why it will not enter the market.
Then we have daily trend remains bearish, but hourly momentum is neutral and price etc. etc. report.html. Let me open it. So, this is how it shows. It shows the the bars, the entry and exit positions with these triangles. I'm going to zoom in in a while, but it also shows the buy and hold. This is the dashed lower curve. That's the buy and hold with the recent gold prices. It also shows the random median of the random agents.
That's the dashed white middle curve. And our agent, the Sol GPT 5.6 is the yellow curve. It looks like it's a bit positive. I mean, it's climbing. It's the equity curve. It's climbing above the random agents for a while, but it's failing to really convince me that it's really guessing a real edge. It's not really having a real edge. It's almost as the best random agent. The cloud, the gray cloud that you can see are all the random agents that we deployed.
So, these are again These are trading agents that are totally flipping a coin for every decision. And this is how it looks like. If you want to analyze further, you can, for example, zoom in into a specific area or a specific slice of the chart. You can see the red triangles are sell positions. Before every triangle, there's a smaller triangle on the previous candle, which tells you there's a sentence that tells you what is the analysis.
For example, short the high volume rebound into hourly EMA 20. Prior break down resist etc. Let me take another one that This is the triangle, which is a short signal. This is the short entry here, but the signal occurred right at the candle before that. And the reason is that re-enter short as the rebound has stalled and rolled over and rolled over below the 20 EMA. So, this is all described here. And actually, what is the agent talking about?
The agent is talking about the downtrend. But then, if you notice at this point, at this level, I don't want to highlight much cuz it's covering the the candles, you have a small rebound up. These four green candles. This is the hourly time frame, by the way. But since it stalled and it got rejected here, and it went back down, at this point, the signal is short. Why? Because there's a small rebound that was rejected in a downtrend.
So, there's no strong retracement. It will continue going down, and this is where the signal was generated. The entry is at the opening price of the following candle. And the take profit was right here. This is the the gray X that you can see at the bottom of this red candle. There's a gray X here. So, this is where we exited, and it's the plus 1.65% take profit. So, it added 1.65% on our on top of our equity. So, did the model find um trading edge, or was it simply lucky?
This is why I created the 20 random agents. They all use the same data, costs, stops, and trading engine. The only difference was that their decisions were random, totally random. And the agent's performance could not be clearly separated from random chance. So, that is the honest result. I know it's kind of boring, but it is what it is. I still have a curious thought, though. What if I change the prompt slightly to push the model to be more aggressive, but I'm not sure if it's going to improve the results by much, knowing how LLMs work.
What I have skipped in this video is the repetition. So, the same experiment should be tested across many different 50-day periods. This was uh only once uh testing this model on 50-days period. So, it's not enough, I guess. Maybe forcing the agent to be more selective might cut fees without uh removing its best signals. And finally, we could test using another model. Could be Claude, for example, and why not Deep Seek or any other model.
And I guess this will be it for today. If you found this experiment useful, you know what the buttons do for this channel. I will see you in the next video. Until our next one, trade safe and see you next time.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.