Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Engineering Visualized · @EngineeringVisualized-1
Words
1,031
Runtime
6:31
Speaking pace
158wpm
Reading time
4min
158 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
What if the architecture behind almost every major AI model today is not the architecture that wins the next decade? Since 2017, transformers have dominated artificial intelligence. Chat, GPT, Claude, Gemini, and most modern language models are built around them. But transformers have weaknesses. Long context becomes expensive. Generation is still mostly one token at a time. Huge models require enormous computation. So researchers are now exploring several different architectures that attack these problems in completely different ways. Some
79 words, the words spoken in the first 30 seconds at 158 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 102 |
| Average words per sentence | 10.1 |
| Longest sentence | 34 words |
| Questions asked | 7 |
| Sentences containing a number | 5 |
Most used terms
Filler phrases
2 in total: like 2.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
What if the architecture behind almost every major AI model today is not the architecture that wins the next decade? Since 2017, transformers have dominated artificial intelligence. Chat, GPT, Claude, Gemini, and most modern language models are built around them. But transformers have weaknesses. Long context becomes expensive. Generation is still mostly one token at a time. Huge models require enormous computation. So researchers are now exploring several different architectures that attack these problems in completely different ways.
Some may replace parts of the transformer. Others may combine with it and one of them could shape what comes next. So let us look at the main challengers. First state space models. The most famous example is Mamba. To understand why it matters, start with a weakness of attention. Imagine a sequence containing 100,000 tokens. A transformer uses attention to compare information across that sequence. That ability is extremely powerful.
But as the sequence grows, full attention becomes increasingly expensive. Mamba approaches the problem differently. Instead of constantly comparing every token with every other token, it carries an internal state through the sequence. Think of reading a book while continuously updating a compact memory of what matters. Mamba also introduced selectivity. It can learn which information should strongly affect that state and which information can mostly be ignored.
The advantage is efficiency. State space models can scale much more gently with sequence length. That makes them attractive for very long context. But there is a problem. Attention can directly jump back to a specific earlier token. A state space model compresses information into its internal state. If an exact detail disappears from that state, retrieving it later can be harder. So Mamba solves the long sequence problem well, but attention can still be better at precise retrieval.
That brings us to challenger number two, modern recurrent models. This sounds strange because recurrent neural networks existed long before transformers. They were supposed to be the old architecture, but they are coming back. Google's recurrent Gemma, for example, uses gated linear recurrence combined with local attention. R W K V follows another recurrent style and XLSTM takes the old LSTM idea and redesigns it for modern large scale models.
The problem recurrence tries to solve is memory efficiency. Instead of keeping a huge attention map, a recurrent model carries information forward through an internal state that can make generation much more memory efficient. But recurrence has its own weakness. Compressing a long history into a state can again make exact retrieval difficult and traditional recurrent models were also harder to train efficiently in parallel.
Modern versions try to solve these problems with better gating, parallel training methods and sometimes small amounts of attention. Notice the pattern already. Attention disappears. Then it comes back. Now challenger number three attacks a completely different problem. Mixture of experts. Suppose a model contains hundreds of billions of parameters. Does every token really need to activate all of them? Mixture of experts says no.
Instead, the model contains many specialist networks called experts. A router examines each token and activates only a small number of them. So, the full model can be enormous while only part of it runs for each token. The problem being solved here is computation. You can increase total model capacity without increasing computation at exactly the same rate. But mixture of experts has an important limitation. It does not really replace the transformer architecture.
Manye models still use attention. It mainly changes which parts of the network activate. So this is less of a transformer replacement and more of a way to make huge models more efficient. Now we reach challenger number four, diffusion language models. This one attacks something completely different. Generation itself. A normal language model generates text like this. One token then the next then the next. Each token depends on what came before it.
Diffusion models try another approach. They can begin with a noisy or incomplete block of text and gradually refine multiple positions. Instead of always writing strictly left to right, the system can work on several parts of the output during the same generation process. Google's diffusion Gemma is a recent example. Google reports that it can generate text significantly faster on suitable hardware because multiple tokens can be produced in parallel.
That sounds like a fundamental alternative. But there is an important twist. Diffusion Gemma still uses a transformer-based backbone. So diffusion is not necessarily replacing transformers either. It is replacing one assumption about them. That text must always be generated one token at a time. And that leads to the fifth and perhaps most important direction. Hybrid architectures. Instead of asking which single architecture wins, researchers are starting to combine them.
AI21's Jamba combines Mamba transformer attention and mixture of experts. IBM's Granite 4 also combines Mamba style layers with transformer attention. Why? Because each architecture is good at something different. State space models are efficient across long sequences. Recurrent systems maintain compact memory. Attention is excellent at directly retrieving specific information. Mixture of experts increases model capacity without activating everything.
Diffusion can change how outputs are generated. A future AI system may use several of these ideas at once. And that may be how transformers are eventually replaced. Not by one architecture suddenly defeating them, but by slowly breaking the transformer apart. Replacing attention where attention is expensive, keeping it where direct retrieval matters, adding recurrence where persistent memory helps, adding experts where capacity matters, and changing generation where token by token decoding becomes the bottleneck.
So what comes after transformers? Right now there is no single winner, but the architecture of modern AI is already becoming less purely transformer-based. And the most interesting possibility is that the model which eventually replaces the transformer may not have one name at all. It may be a combination of everything that worked. What do you think? Do you think transformers will still dominate AI 10 years from now?
Or are we already seeing the architecture that replaces them being assembled piece by piece? Let me know in the comments. This is Engineering Visualize. Subscribe for more visual explanations of the architectures and algorithms behind modern AI.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.