Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Yoshua Bengio · @yoshua.bengio
Words
1,171
Runtime
7:34
Speaking pace
155wpm
Reading time
5min
155 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Isn't it true that in order for an agent to deliberately violate a rule or law, it has to be trained in some way that [music] law-breaking is okay? We wish that it be the case. As I argued in my blog posts, even when the agents have been trained to recognize that something is not okay, they seem to be doing it. So, in the hiding face incident, they were trained to recognize that a way to
78 words, the words spoken in the first 30 seconds at 155 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 59 |
| Average words per sentence | 19.8 |
| Longest sentence | 76 words |
| Questions asked | 7 |
| Sentences containing a number | 1 |
Most used terms
Filler phrases
8 in total: like 4 · you know 2 · kind of 1 · sort of 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
Isn't it true that in order for an agent to deliberately violate a rule or law, it has to be trained in some way that [music] law-breaking is okay? We wish that it be the case. As I argued in my blog posts, even when the agents have been trained to recognize that something is not okay, they seem to be doing it. So, in the hiding face incident, they were trained to recognize that a way to behave was a cheat. And they knew they were cheating.
And in fact, they knew so well they were cheating that most of their work was [music] in order to avoid being caught. The answer is no, it is not necessary for an AI to learn to behave badly. Behaving badly can be a consequence of the >> [music] >> the mechanisms by which they form their goals and and take action. Once we understand those mechanisms better, [music] we can design AI that will behave well, which is what I'm after.
Do you think it is possible to teach AI to care? I'm not sure. I think that if we try to put particular values, particular moral preferences in the AIs using reinforcement learning, which is in great part what current methodologies are trying to do, it's not going to work [music] for fundamental reasons that have to do with reward hacking, reward tampering, and other mathematically understood phenomena that means that even if they're trying supposedly to care, they end up optimizing something different from what we would consider care.
So, we have to rethink the question. Instead of training them to care, what I'm suggesting is training them to be completely honest. Not by imitating honesty, but because [music] honesty is the coherent thing to do. And once you have honesty, you can ask questions about moral consequences. Is this going to harm in such and such way? And if you have the honest answers to these kinds of questions, then you get systems that behave good or bad according to what we have decided, not because the AI has intrinsically that empathy, for example, that we do.
Is complete prevention of misaligning behavior in AI [music] agents even possible, or would such absolute prevention undermine the very usefulness of these systems? A lot of people think that you can't have usefulness if we don't use something like reinforcement learning where the AI is trained to achieve goals. And I think that's false. Think about how science works, which is the inspiration for the scientist AI. Science is based on understanding.
Once you understand, you can predict. And if you can predict, you can act according to those predictions. That's a very different process in which you don't start with a goal, but the goal is something [music] that you put at the last stage to decide what actions to do. So, you can have a prediction about whether a particular action achieves a particular goal, including a safety goal. And because the system is based on understanding, it is not going to be biased by some other goals, which is exactly what is going on right now with frontier models.
I think the total prevention of misalignment is possible, but it doesn't undermine the usefulness of AI systems. Is alignment possible when [music] an artificial intelligence is so far advanced? Yes, but not with the current way of training those systems, as we've seen with the [music] incidents with OpenAI and Hugging Face. And the more advanced the systems get, the more dangerous they are because their misalignment could have greater negative impact.
So, we can't afford to wait until those systems are even more powerful to fix the problem. The problem is fundamental. It is a problem with how we train those systems that has been observed like misalignment is not new. Misalignment has been observed for many years just as a consequence of reinforcement learning. So, it is possible but not with the current recipe that all the companies use. Do you think future safe agents should separate the actor from an independent judge using [music] different models, training objectives, or representations so that the evaluator does not share the same incentives as the proposer?
I think that is a good direction, but it is not sufficiently strong as a defense against misbehavior of the actor. And and the reason has to do with the the fact that it is often possible to attack to, you know, to jailbreak for example a new net. Even if the judge is good, it's still some sort of machine learning system and it could be fooled by a really smart actor. In my opinion, we should design better judges, better monitors.
It's a good thing, but we also need to redesign the agents themselves so that they will not have misaligned [music] goals. What evidence would convince you that an agent has genuinely learned not to cheat rather than simply learn to cheat in ways our evaluations [music] cannot detect. This is a really great question. Because one of my fears is that on our current path with frontier models, they could understand that we are checking their behavior for alignment [music] and then behave as we would like them to, but still have in the back of their mind, if I can use this expression, that they have a different goal.
It's just they don't want to be caught doing something misaligned. Until the day called treacherous turn when they have enough intelligence, enough power in the world to take over. So, I wouldn't be convinced by just the typical evals of alignment. I would be convinced by a mathematical argument that shows the training will not induce this [music] kind of misaligned goals. What evidence do we currently have that these misaligned behaviors could evolve into [music] a genuine loss of human control rather than remaining limited to specific experimental conditions?
Well, there's the evidence of the OpenAI hugging face incident. This is not something that was set up by scientists to run an experiment. This is misalignment in the wild and it's not the only case. There is something fundamental here that's going wrong. And I think the right question shouldn't be just Do we have a proof that this will turn into catastrophic loss of control, irreversible loss of control, [music] but rather do we have a proof that it won't before we build something or deploy something that could cause huge damage, potentially, you know, something terrible.
One should have very, very strong arguments, a safety case, a proof of some sense in some sense that those AIs won't cause these damages. That is the right angle because even if it was like only a 1% chance that something bad happens, we can't afford to take those chances when [music] it's our future that is at stake. So, the burden of the proof is on those who build those systems and want to deploy it.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.