| I BUILT A FULLY AUTOMATIC MANSPLAINER | Mar 6, 2026 | 13:46 | 12.5K |
| Traditional X-Mas Stream | Dec 29, 2025 | 2:33:37 | 4.5K |
| Traditional Holiday Live StreamShort | Dec 28, 2025 | 0:00 | 0 |
| TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis) | Dec 27, 2025 | 47:02 | 24.2K |
| Titans: Learning to Memorize at Test Time (Paper Analysis) | Dec 14, 2025 | 32:31 | 24.9K |
| [Paper Analysis] The Free Transformer (and some Variational Autoencoder stuff) | Nov 1, 2025 | 40:10 | 23.7K |
| [Video Response] What Cloudflare's code mode misses about MCP and tool calling | Oct 19, 2025 | 13:19 | 9.8K |
| [Paper Analysis] On the Theoretical Limitations of Embedding-Based Retrieval (Warning: Rant) | Oct 11, 2025 | 48:57 | 37K |
| AGI is not coming! | Aug 9, 2025 | 7:09 | 161K |
| Context Rot: How Increasing Input Tokens Impacts LLM Performance (Paper Analysis) | Jul 23, 2025 | 37:49 | 33.7K |
| Energy-Based Transformers are Scalable Learners and Thinkers (Paper Review) | Jul 19, 2025 | 47:51 | 32.8K |
| On the Biology of a Large Language Model (Part 2) | May 3, 2025 | 56:26 | 17.7K |
| On the Biology of a Large Language Model (Part 1) | Apr 5, 2025 | 54:05 | 59.4K |
| [GRPO Explained] DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models | Jan 26, 2025 | 1:09:00 | 175.1K |
| Traditional Holiday Live Stream | Dec 27, 2024 | 1:28:17 | 5.2K |
| Byte Latent Transformer: Patches Scale Better Than Tokens (Paper Explained) | Dec 24, 2024 | 36:15 | 48.2K |
| Safety Alignment Should be Made More Than Just a Few Tokens Deep (Paper Explained) | Dec 10, 2024 | 48:53 | 13.6K |
| TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters (Paper Explained) | Nov 23, 2024 | 28:23 | 19.8K |
| GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models | Oct 19, 2024 | 37:06 | 21.2K |
| Were RNNs All We Needed? (Paper Explained) | Oct 12, 2024 | 27:48 | 60.6K |
| Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters (Paper) | Oct 5, 2024 | 53:02 | 31.6K |
| Privacy Backdoors: Stealing Data with Corrupted Pretrained Models (Paper Explained) | Aug 4, 2024 | 1:03:56 | 19K |
| Scalable MatMul-free Language Modeling (Paper Explained) | Jul 8, 2024 | 49:45 | 35.4K |
| Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools (Paper Explained) | Jun 26, 2024 | 1:11:58 | 42.3K |
| xLSTM: Extended Long Short-Term Memory | Jun 1, 2024 | 57:00 | 45.8K |
| [ML News] OpenAI is in hot waters (GPT-4o, Ilya Leaving, Scarlett Johansson legal action) | May 21, 2024 | 29:22 | 33.6K |
| ORPO: Monolithic Preference Optimization without Reference Model (Paper Explained) | May 1, 2024 | 33:26 | 26.1K |
| [ML News] Chips, Robots, and Models | Apr 30, 2024 | 39:14 | 29.4K |
| TransformerFAM: Feedback attention is working memory | Apr 28, 2024 | 37:01 | 39.9K |
| [ML News] Devin exposed | NeurIPS track for high school students | Apr 27, 2024 | 17:47 | 41.2K |
| Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention | Apr 24, 2024 | 37:17 | 60.4K |
| [ML News] Llama 3 changes the game | Apr 23, 2024 | 31:19 | 47.9K |
| Hugging Face got hacked | Apr 17, 2024 | 18:01 | 31.7K |
| [ML News] Microsoft to spend 100 BILLION DOLLARS on supercomputer (& more industry news) | Apr 15, 2024 | 9:55 | 21.9K |
| [ML News] Jamba, CMD-R+, and other new models (yes, I know this is like a week behind 馃檭) | Apr 13, 2024 | 27:32 | 25.8K |
| Flow Matching for Generative Modeling (Paper Explained) | Apr 8, 2024 | 56:16 | 110.9K |
| Beyond A*: Better Planning with Transformers via Search Dynamics Bootstrapping (Searchformer) | Apr 6, 2024 | 44:05 | 37.1K |
| [ML News] Grok-1 open-sourced | Nvidia GTC | OpenAI leaks model names | AI Act | Mar 26, 2024 | 27:00 | 34.1K |
| [ML News] Devin AI Software Engineer | GPT-4.5-Turbo LEAKED | US Gov't Report: Total Extinction | Mar 17, 2024 | 26:50 | 52.2K |
| [ML News] Elon sues OpenAI | Mistral Large | More Gemini Drama | Mar 10, 2024 | 53:15 | 32.3K |
| On Claude 3Short | Mar 7, 2024 | 1:00 | 11.8K |
| No, Anthropic's Claude 3 is NOT sentient | Mar 5, 2024 | 15:12 | 44.3K |
| [ML News] Groq, Gemma, Sora, Gemini, and Air Canada's chatbot troubles | Mar 1, 2024 | 42:34 | 42.7K |
| Gemini has a Diversity Problem | Feb 22, 2024 | 17:36 | 54.4K |
| V-JEPA: Revisiting Feature Prediction for Learning Visual Representations from Video (Explained) | Feb 19, 2024 | 50:03 | 63.1K |
| What a day in AI! (Sora, Gemini 1.5, V-JEPA, and lots of news) | Feb 18, 2024 | 1:23:59 | 32.7K |
| Lumiere: A Space-Time Diffusion Model for Video Generation (Paper Explained) | Feb 4, 2024 | 54:24 | 30.7K |
| AlphaGeometry: Solving olympiad geometry without human demonstrations (Paper Explained) | Jan 21, 2024 | 35:27 | 40.6K |
| Mixtral of Experts (Paper Explained) | Jan 13, 2024 | 34:32 | 66.9K |
| Until the Litter End | Jan 10, 2024 | 3:40 | 13.8K |