YouTube transcripts

AI Might Be Conscious, But Not As We Thought: video thumbnail

AI Might Be Conscious, But Not As We Thought transcript

Sabine Hossenfelder · @SabineHossenfelder

Published September 22, 20266:2998.1K views

Watch this video on YouTube

Transcript analysisComputed from the caption text

Words

1,015

Runtime

6:29

Speaking pace

157wpm

Reading time

4min

157 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.

Opening (first 30 seconds)

I'm convinced that artificial intelligence will become conscious eventually. There have now been new rumors again that the current AI might be conscious already until little notice paper that basically argues that they may be conscious not while you use them but during the prior phase of training. So let's have a look. Claude can watch its own thoughts, but connected to the right tools. It can do something more useful. It can make your thoughts turn into reality.

79 words, the words spoken in the first 30 seconds at 157 words per minute.

Sentence shape

MeasureThis transcript
Sentences70
Average words per sentence14.5
Longest sentence40 words
Questions asked2
Sentences containing a number6

Most used terms

  • claude17
  • conscious12
  • models9
  • input7
  • language7
  • language models7
  • large7
  • large language7
  • anthropic6
  • consciousness5
  • hoel5
  • internal5

Filler phrases

6 in total: basically 2 · like 2 · actually 1 · sort of 1.

A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.

What this transcript is

Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, published by the channel, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.

Transcript

I'm convinced that artificial intelligence will become conscious eventually. There have now been new rumors again that the current AI might be conscious already until little notice paper that basically argues that they may be conscious not while you use them but during the prior phase of training. So let's have a look. Claude can watch its own thoughts, but connected to the right tools. It can do something more useful.

It can make your thoughts turn into reality. Our sponsor Higgsfield just released Seedance 2.5 and you can connect it to Claude. I tried it. It really only takes about a minute. You just add it to the Claude connectors. Here we give it one prompt. Direct a 30-second film in one take about a physicist who finds a thought in her head that isn't hers. Claude plant the shots and Seedance 2.5 generated what you're watching right now.

It does this in one take or 30 seconds without awkward cuts. Same face, same light, same room. And if a detail is wrong, you can fix that one detail without having to regenerate the entire thing. This used to be the part where AI videos fell apart, but Seedance 2.5 makes it work. Links in the description below. So, go and direct something. Thanks for watching. See you tomorrow. And now back to the science news. Claude is the chatbot made by Anthropic and it’s certainly the one that has attracted the most chatter.

Richard Dawkins, author of the “God Delusion” claimed in an essay in May that he can’t rule out that Claude is conscious. His argument is basically that he thinks Claude is so good that if this isn’t consciousness, then what do we even need consciousness for which, if nothing else, tells us that Dawkins hasn’t spent much time with Large Language Models. But Anthropic’s CEO Dario Amodei himself has said “We don’t know if the models are conscious… But we’re open to the idea that it could be.” It’s not just words.

In a new paper that just appeared, Anthropic researchers say they found a small, special part of Claude’s internal activity that works like a silent scratchpad. Claude can put information there and use it for reasoning, without ever outputting it. This is very interesting, because it means that Claude has a sort of internal life and is capable of introspection. It’s one of the necessary requirements for global workspace theory, that is one of the leading theories for consciousness.

More formally, the Anthropic team called the internal workspace the J-space, where the J stands for Jacobian matrix, that’s the mathematical object that the researchers calculate to identify the space. The Anthropic researchers now showed that if they change what is in Claude’s J-space, Claude’s answer changes. For example if Claude has internally worked out that the animal which spins webs is a spider, and the researchers replace “spider” with “ant,” Claude answers as if the animal had six legs, not eight.

One Anthropic researcher, Jack Lindsey, already claimed he had evidence for Claude’s introspection last year. The new paper now is a fuller analysis that also identifies exactly where the introspection happens, namely in said J-spaces. It’s not just Claude. Another group likewise injected words into Llama’s internal activity and found that Llama also seemed to take note of it. Probably a similar thing is going on with all Large Language Models.

But not everyone is convinced. Researchers from New York University have challenged these claims. They tested this with three LLMs and found that the models could not reliably distinguish a change that was made to their internal activity from an ordinary prompt. So they are saying it’s just yet another way to give them input. Then we have Erik Hoel, a neuroscientist who works on consciousness. He has written an interesting paper in which he argues that today’s large language models are probably not conscious.

We’ve heard this before, of course, but his argument is new. Hoel compares Large Language Models to lookup tables. Such “lookup table” algorithms are the classical examples that demonstrate that observing input and output alone cannot tell you whether a system is conscious. It’s like looking up Chinese translations doesn’t mean you understand Chinese. Hoel now argues that by way of computational structure, Large Language Models are much closer to lookup tables than they are to the human brain.

They are stochastic input-output machines. Much of the autonomy that we assign to conscious beings on the other hand is not directly a reaction to input. And Hoel says that the most obvious reason why LLMs are not conscious is that they can’t learn from new input, though that has not stopped quite a few people from becoming senior management. This leads to a strange possibility. You see, large language models work in two stages.

The first is the training in which they get fed a lot of input. Once this is done, you freeze the fully trained LLM and use it as an input output machine. These are the apps you sign up for. Yes, they now use memory, but this is externally latched on as context to the prompt, it does not actually update the model. If Hoel is right, then maybe large language models are not conscious when we use them. But maybe they are conscious during training.

Personally I don’t think this is a relevant distinction. Imagine you have a human being unable to form new memories, you wouldn’t conclude that therefore they are unconscious, would you? That said, to me the question of whether an LLM is conscious or not makes no sense, the question is how we quantify its consciousness. This is why I give all these papers a 9 out of 10 on the bullshit meter. They fail at the relevant question: if you can’t measure it, what are you even talking about?

Though I have to admit if Claude has thoughts it does not say out loud. This already puts it ahead of most people on social media. Thanks for watching. See you tomorrow.

The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.

Use this transcript

Three free tools that work on the material around a video like this one. No signup, no login.