<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The Information Bottleneck]]></title><description><![CDATA[AI research, compressed. Long conversations with the people building the frontier, and a working researcher's take on the ideas that survive the bottleneck - minus the hype.]]></description><link>https://www.the-information-bottleneck.com</link><image><url>https://substackcdn.com/image/fetch/$s_!nQnk!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9b10938-f656-4406-aa7a-36b5e263a5dc_950x950.png</url><title>The Information Bottleneck</title><link>https://www.the-information-bottleneck.com</link></image><generator>Substack</generator><lastBuildDate>Mon, 28 Sep 2026 11:18:34 GMT</lastBuildDate><atom:link href="https://www.the-information-bottleneck.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[The Information Bottleneck]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[informationbottleneck@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[informationbottleneck@substack.com]]></itunes:email><itunes:name><![CDATA[Ravid Shwartz Ziv]]></itunes:name></itunes:owner><itunes:author><![CDATA[Ravid Shwartz Ziv]]></itunes:author><googleplay:owner><![CDATA[informationbottleneck@substack.com]]></googleplay:owner><googleplay:email><![CDATA[informationbottleneck@substack.com]]></googleplay:email><googleplay:author><![CDATA[Ravid Shwartz Ziv]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Why AI Still Can't See wtih Andrew Dai]]></title><description><![CDATA[Listen now (58 mins) | Andrew Dai spent over a decade at Google Brain and DeepMind, where he co-wrote the 2015 paper that introduced language model pre-training followed by fine-tuning, and later co-led pre-training data for Gemini.]]></description><link>https://www.the-information-bottleneck.com/p/why-ai-still-cant-see-wtih-andrew-1ef</link><guid isPermaLink="false">https://www.the-information-bottleneck.com/p/why-ai-still-cant-see-wtih-andrew-1ef</guid><dc:creator><![CDATA[Ravid Shwartz Ziv]]></dc:creator><pubDate>Thu, 24 Sep 2026 20:08:00 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/217297838/371080e683149fc0f88af5712bf68135.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-p2VPLwIvM1Y" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;p2VPLwIvM1Y&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/p2VPLwIvM1Y?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>Andrew Dai spent over a decade at Google Brain and DeepMind, where he co-wrote the 2015 paper that introduced language model pre-training followed by fine-tuning, and later co-led pre-training data for Gemini. He's now co-founder and CEO of Elorian AI, which is building models for visual reasoning.</p><p>We talk about how his pre-training result started as a bug, why next-token prediction scales better than other objectives, and what went wrong for Google in the early LLM race. The second half is about vision: why today's frontier models still can't count objects in a photo, why he thinks reasoning is fundamentally visual, and how his view of world models differs from JEPA.</p><p><strong>Chapters</strong></p><p>00:00 Intro<br>00:48 Andrew's background and the accidental discovery of pre-training<br>07:10 Why next-token prediction scales<br>10:09 How Google fell behind and the early days of Gemini<br>18:28 What makes training data good<br>28:16 Where visual understanding breaks down<br>46:01 World models, JEPA and robotics<br>55:27 Generation vs understanding, and what's next for Elorian</p><p><strong>Topics</strong></p><ul><li><p>Pre-training and fine-tuning</p></li><li><p>Scaling and next-token prediction</p></li><li><p>Gemini and Google's AI history</p></li><li><p>Data quality and synthetic data</p></li><li><p>Visual reasoning and counting</p></li><li><p>World models and JEPA</p></li></ul><p><strong>Music</strong></p><ul><li><p>"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0</p></li></ul>]]></content:encoded></item><item><title><![CDATA[Can AI Replace Doctors? Zachary Lipton on the Future of Healthcare and AI]]></title><description><![CDATA[Zachary Lipton is an associate professor at Carnegie Mellon University and a co-founder of Abridge, a healthcare AI company (and a jazz saxophonist!).]]></description><link>https://www.the-information-bottleneck.com/p/can-ai-replace-doctors-zachary-lipton-be1</link><guid isPermaLink="false">https://www.the-information-bottleneck.com/p/can-ai-replace-doctors-zachary-lipton-be1</guid><dc:creator><![CDATA[Ravid Shwartz Ziv]]></dc:creator><pubDate>Mon, 21 Sep 2026 04:18:53 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/216680707/bdd28324039bd16265367de7565390eb.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-RhRL9dZ7A1Q" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;RhRL9dZ7A1Q&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/RhRL9dZ7A1Q?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>Zachary Lipton is an associate professor at Carnegie Mellon University and a co-founder of Abridge, a healthcare AI company (and a jazz saxophonist!). His research spans machine learning, healthcare, and the broader impact of AI.</p><p>In this episode, we discuss what AI can (and cannot yet) do in healthcare. We talk about why medicine is harder to automate than coding, AI scribes and clinical decision support, drug discovery, whether AI could eventually replace doctors, and the role of open models and intelligent routing.</p><p>In the second half, we turn to the future of AI research: whether academia can still compete, what AI PhD students should work on, and how Zach thinks about automation and the future of research.</p><h3>Timeline</h3><p>00:00 &#8212; Introduction<br>00:28 &#8212; Why healthcare is hard for AI<br>05:00 &#8212; Zach&#8217;s path into healthcare AI<br>22:13 &#8212; Where AI can have the biggest impact<br>32:01 &#8212; Can AI replace doctors?<br>41:41 &#8212; Open models and AI infrastructure<br>56:20 &#8212; The future of AI research<br>1:03:03 &#8212; What should AI PhD students work on?<br>1:27:00 &#8212; Advice for researchers and founders</p><h3>Topics</h3><p>AI in healthcare &#8226; AI doctors &#8226; drug discovery &#8226; clinical decision support &#8226; open-source AI &#8226; AI research &#8226; academia vs. industry &#8226; future of work</p><p><strong>Music</strong></p><ul><li><p>"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0</p></li></ul>]]></content:encoded></item><item><title><![CDATA[Sara Hooker on the End of Static AI]]></title><description><![CDATA[What comes after scaling?]]></description><link>https://www.the-information-bottleneck.com/p/sara-hooker-on-the-end-of-static-2fa</link><guid isPermaLink="false">https://www.the-information-bottleneck.com/p/sara-hooker-on-the-end-of-static-2fa</guid><dc:creator><![CDATA[Ravid Shwartz Ziv]]></dc:creator><pubDate>Wed, 16 Sep 2026 21:10:16 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/216060454/4fea170ea010ee72b3617dd27e1bea72.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p></p><div id="youtube2-O_ZRN5UI5BU" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;O_ZRN5UI5BU&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/O_ZRN5UI5BU?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>We talk with <strong>Sara Hooker, co-founder and CEO of Adaptation Lab</strong>, about why the next generation of AI may look very different from today's static models. Sara argues that models should continuously adapt to new tasks, data, users, and environments&#8212;and that doing this efficiently will require rethinking much more than fine-tuning.</p><p>We discuss continual learning, AutoScientist and automated research, why non-verifiable tasks may become the next major bottleneck, and why interfaces could be as important as the models themselves. We also get into open vs. closed models, distillation and Chinese AI labs, AI regulation and safety, cybersecurity and biorisk, AI companionship, and what may eventually come after Transformers and tokenization.</p><h3>Topics</h3><ul><li><p>Continuous learning and adaptive AI</p></li><li><p>Fine-tuning, memory, and AutoScientist</p></li><li><p>AI agents and automated research</p></li><li><p>Non-verifiable tasks and human feedback</p></li><li><p>Adaptive interfaces</p></li><li><p>Open vs. closed models and distillation</p></li><li><p>AI safety, regulation, cyber risk, and biorisk</p></li><li><p>AI companionship and persuasion</p></li><li><p>The limits of Transformers</p></li><li><p>Multilingual models and tokenization</p></li></ul><h3>Chapters</h3><p><strong>00:00</strong> &#8212; Introduction<br><strong>02:15</strong> &#8212; Why start another AI lab? The return of research<br><strong>05:46</strong> &#8212; What continuous learning actually means<br><strong>12:04</strong> &#8212; Should every company have its own adapting model?<br><strong>13:59</strong> &#8212; Fine-tuning and platforms like Tinker<br><strong>18:04</strong> &#8212; AutoScientist and automated optimization<br><strong>22:52</strong> &#8212; Can AI really improve its own research?<br><strong>28:38</strong> &#8212; The problem of non-verifiable tasks<br><strong>31:30</strong> &#8212; Human feedback and the limits of exponential progress<br><strong>34:43</strong> &#8212; Why the AI interface matters<br><strong>40:36</strong> &#8212; Distillation, China, and open models<br><strong>49:05</strong> &#8212; Open-model licensing<br><strong>52:19</strong> &#8212; Will open models catch closed models?<br><strong>58:43</strong> &#8212; AI regulation and compute thresholds<br><strong>1:03:07</strong> &#8212; AI safety and agent failures<br><strong>1:10:19</strong> &#8212; Biorisk vs. cybersecurity<br><strong>1:14:03</strong> &#8212; Persuasion, AI companionship, and overlooked risks<br><strong>1:20:41</strong> &#8212; Where will AI have the biggest real-world impact?<br><strong>1:25:41</strong> &#8212; What is missing from current AI architectures?<br><strong>1:29:03</strong> &#8212; Neurosymbolic AI<br><strong>1:31:30</strong> &#8212; Multilingual models and tokenization<br><strong>1:34:02</strong> &#8212; Byte-level models and alternatives to tokenization<br><strong>1:35:03</strong> &#8212; Closing</p><p><strong>Music</strong></p><ul><li><p>"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0</p></li></ul>]]></content:encoded></item><item><title><![CDATA[Tiny Recursive Models Beat the Giants - Alexia Jolicoeur-Martineau (Microsoft)]]></title><description><![CDATA[Alexia Jolicoeur-Martineau is a Principal Researcher at Microsoft and the author of "Less is More: Recursive Reasoning with Tiny Networks," the paper behind the Tiny Recursive Model that hit about 45% on ARC-AGI-1 with a fraction of the parameters of frontier systems.]]></description><link>https://www.the-information-bottleneck.com/p/tiny-recursive-models-beat-the-giants-a55</link><guid isPermaLink="false">https://www.the-information-bottleneck.com/p/tiny-recursive-models-beat-the-giants-a55</guid><dc:creator><![CDATA[Ravid Shwartz Ziv]]></dc:creator><pubDate>Tue, 15 Sep 2026 04:04:37 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/215769336/9422d52415fb58db69626845b824f5d4.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-lTTs4W3F7jg" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;lTTs4W3F7jg&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/lTTs4W3F7jg?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>Alexia Jolicoeur-Martineau is a Principal Researcher at Microsoft and the author of "Less is More: Recursive Reasoning with Tiny Networks," the paper behind the Tiny Recursive Model that hit about 45% on ARC-AGI-1 with a fraction of the parameters of frontier systems. It won the 2025 ARC Prize paper award.</p><p>She read the hierarchical reasoning paper, thought the potential was real and the explanation was not, and rebuilt it without the mouse brains: a small network that carries a hidden state and a current answer, thinks for a few steps, updates, and repeats, with the gradient truncated at each loop. We get into why puzzles suit this and autoregression doesn't, why she thinks LLMs are bad at molecules and more data won't fix it, and what she'd do with a trillion dollars.</p><div><hr></div><p><strong>Timeline</strong></p><ul><li><p>00:01 Intro</p></li><li><p>01:06 Leaving biostatistics, and why the field stagnated</p></li><li><p>06:47 GANs, diffusion, and research on four GPUs</p></li><li><p>12:58 What was wrong with the hierarchical reasoning paper</p></li><li><p>16:31 Tiny recursive models explained without the biology</p></li><li><p>22:35 Why puzzles favor recursion over left to right generation</p></li><li><p>24:15 Is the bitter lesson really bitter?</p></li><li><p>27:28 With infinite compute, would you still want small models?</p></li><li><p>32:00 Self improvement, memory, and a trillion dollars</p></li><li><p>37:01 Test time compute beyond chain of thought</p></li><li><p>40:41 Why chain of thought fails on molecules</p></li><li><p>45:17 Is there a universal representation?</p></li><li><p>48:06 What people are already building with TRM</p></li><li><p>55:22 Fixed point models and DEQ</p></li><li><div><hr></div><p><strong>Music</strong></p></li></ul><ul><li><p>"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0</p></li></ul><p><strong>Topics</strong></p><ul><li><p>Tiny Recursive Models and the ARC-AGI results</p></li><li><p>What the hierarchical reasoning model was really doing</p></li><li><p>Deep supervision and truncated backprop</p></li><li><p>Looping transformers and parameter efficiency</p></li><li><p>Why puzzles favor whole-context iteration over left to right generation</p></li><li><p>Test time compute beyond chain of thought</p></li><li><p>Latent reasoning and the Coconut line of work</p></li><li><p>Why LLMs fail on chemistry and physics</p></li><li><p>Representation learning and whether a universal representation exists</p></li></ul><h1></h1>]]></content:encoded></item><item><title><![CDATA[Continual Learning Is the Next Bottleneck | Rohan Anil (Core Automation )]]></title><description><![CDATA[Listen now (71 mins) | Rohan Anil spent eleven and a half years at Google, where he went from writing memory allocators to large-scale linear solvers, then optimization at Google Brain, where he co-developed distributed Shampoo and led optimization for PaLM and Gemini pre-training, including the work that produced Gemini Flash.]]></description><link>https://www.the-information-bottleneck.com/p/continual-learning-is-the-next-bottleneck-2fc</link><guid isPermaLink="false">https://www.the-information-bottleneck.com/p/continual-learning-is-the-next-bottleneck-2fc</guid><dc:creator><![CDATA[Ravid Shwartz Ziv]]></dc:creator><pubDate>Thu, 10 Sep 2026 03:40:22 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/214987237/0512cad7f93087eafb636d45a4767988.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-AQZF_jJeMJs" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;AQZF_jJeMJs&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/AQZF_jJeMJs?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>Rohan Anil spent eleven and a half years at Google, where he went from writing memory allocators to large-scale linear solvers, then optimization at Google Brain, where he co-developed distributed Shampoo and led optimization for PaLM and Gemini pre-training, including the work that produced Gemini Flash. He then joined Anthropic's pre-training team, and left before the IPO to co-found Core Automation with Jerry Tworek (ex-VP of Research at OpenAI). We talk with him about how Brain worked at its peak, why he left two of the world's best labs, and what he thinks is missing from today's models.</p><p>Rohan's view is that pre-training and RL were split by organizational convenience rather than by science. Pre-training builds a prior, and RL sharpens it to the tasks we care about, and neither gives a model a way to absorb new data or learn from its own experience once it is deployed. Post-training more every day plateaus, on-policy distillation plateaus, and in-context learning only goes as far as the context does. He argues the next architecture needs better ways to fold in new knowledge at inference time, and that this is a fundamental optimization question rather than a harness-engineering one.</p><p>We also get into why coding agents still fail on low-level systems work, his take on Muon, why second-order methods matter once you leave the noise-dominated regime, and why nobody can yet use a few million GPUs for a single training run.</p><div><hr></div><p><strong>Timeline</strong></p><p>00:00 Intro<br>01:09 From computer vision to Google systems engineering<br>02:37 Large-scale linear solvers and sparse features<br>06:01 Getting into optimization: SDCA and Yonghui Wu's team<br>07:29 Joining the Shampoo crew<br>09:35 The Google Brain ethos, and why 2017 to 2019 was special<br>14:23 Is open research going to keep winning?<br>15:53 Frontier models are only as good as the prior you give them<br>17:30 Missing the language model wave, then Common Crawl and online distillation<br>18:31 Paternity leave, DALL-E Mini, and the 14 days that became two years<br>20:32 PaLM, Gemini pre-training, and Gemini Flash<br>23:59 The Shampoo origin story: Tomer Koren's two-week proof<br>26:54 Why leave Google for Anthropic<br>30:20 Why leave Anthropic for a startup<br>31:33 Meeting Jerry Tworek at Dolores Park<br>34:00 What Core Automation is building<br>36:17 Continual learning and the pre-training vs RL split<br>39:11 Why coding agents fail at kernels and low-level pipelines<br>42:00 The QR factorization kernel competition and reward hacking<br>44:57 Numerics, verification, and hardware that keeps changing<br>46:07 Are LLMs creative, or just good at search?<br>49:50 Getting models to extrapolate instead of interpolate<br>52:03 Why did we ever call it pre-training?<br>55:02 What RL is really learning<br>56:27 Competing with the big labs with fewer people<br>58:31 Will kernel generation keep old GPUs alive? Amdahl's law<br>1:01:58 Open source plans<br>1:02:55 Audience question: agentic optimizers<br>1:04:53 Audience question: Muon, Shampoo, and the future of second-order methods<br>1:08:58 Hiring at Core Automation</p><div><hr></div><p><strong>key topics</strong></p><ul><li><p>Journey from Google Brain to startup</p></li><li><p>Evolution of AI research and optimization</p></li><li><p>Pre-training and reinforcement learning</p></li><li><p>Kernel optimization and system efficiency</p></li><li><p>Open source AI and collaborative research</p></li><li><p>Challenges in AI creativity and exploration</p></li><li><p>Future directions in continual learning and model scaling</p><div><hr></div></li></ul><ul><li><p><strong>Music</strong></p></li></ul><ul><li><p>"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0</p></li></ul>]]></content:encoded></item><item><title><![CDATA[World Models | John Langford (Microsoft AI Labs)]]></title><description><![CDATA[John Langford, one of the heads of Microsoft's AI Labs, the creator of Vowpal Wabbit, and a co-inventor of CAPTCHA, joins us to talk about world models.]]></description><link>https://www.the-information-bottleneck.com/p/world-models-john-langford-microsoft-48c</link><guid isPermaLink="false">https://www.the-information-bottleneck.com/p/world-models-john-langford-microsoft-48c</guid><dc:creator><![CDATA[Ravid Shwartz Ziv]]></dc:creator><pubDate>Sat, 05 Sep 2026 00:32:46 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/214241039/c8c3e7fdc9b1e470d3c05755103f6ee1.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-0_cAVd8rTw0" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;0_cAVd8rTw0&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/0_cAVd8rTw0?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>John Langford, one of the heads of Microsoft's AI Labs, the creator of Vowpal Wabbit, and a co-inventor of CAPTCHA, joins us to talk about world models. Transformers need orders of magnitude more data than humans to learn the same thing, and John argues a compact, implicit world model is how you close that gap. He explains why he's skeptical of JEPA-style objectives, why a transformer's KV cache is the Ptolemaic epicycle model of belief states, and what his Next Latent work does differently.</p><p>We also get into whether research still matters in the age of scale; open versus closed models; agent-driven research after running 2,000 pre-training experiments in 90 days; the origin story of CAPTCHA; and why Muon and orthonormal optimizers actually work.</p><div><hr></div><p><strong>Topics:</strong></p><ul><li><p>Implicit vs. explicit world models, and the case against JEPA-style objectives</p></li><li><p>Compact belief states: why compression beats a growing KV cache</p></li><li><p>Does research still matter in the age of scale? The Kimi K3 argument</p></li><li><p>Agent-driven research: 2,000 pre-training experiments in 90 days</p></li><li><p>The invention of CAPTCHA</p></li><li><p>Optimizers from SGD and Vowpal Wabbit to Muon</p></li><li><div><hr></div></li></ul><h2>Chapters</h2><ul><li><p>00:00 Why world models: the sample-complexity gap</p></li><li><p>09:48 The case against JEPA; a transformer-style implicit world model</p></li><li><p>15:52 Compact belief states: epicycles vs. heliocentrism</p></li><li><p>23:41 Does research still matter? The Kimi K3 argument</p></li><li><p>27:35 Open vs. closed models</p></li><li><p>35:57 Recursive self-improvement and agent-driven research</p></li><li><p>42:30 2,000 pre-training experiments in 90 days; weak baselines and reproducibility</p></li><li><p>54:54 The invention of CAPTCHA</p></li><li><p>1:00:53 Optimizers: from Vowpal Wabbit to Muon</p></li><li><div><hr></div><p><strong>Music</strong></p></li></ul><ul><li><p>"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0</p></li></ul>]]></content:encoded></item><item><title><![CDATA[Which Tabular Model Should You Actually Use? | David Holzmüller (INRIA)]]></title><description><![CDATA[Tabular data is still where most of machine learning actually happens in industry, and the field has changed a lot in the last few years.]]></description><link>https://www.the-information-bottleneck.com/p/which-tabular-model-should-you-actually-e82</link><guid isPermaLink="false">https://www.the-information-bottleneck.com/p/which-tabular-model-should-you-actually-e82</guid><dc:creator><![CDATA[Ravid Shwartz Ziv]]></dc:creator><pubDate>Thu, 03 Sep 2026 14:30:06 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/214019410/984501987a4b8c4fa4303afa7c01cb87.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p></p><div id="youtube2-T9H0TdOxZo4" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;T9H0TdOxZo4&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/T9H0TdOxZo4?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>Tabular data is still where most of machine learning actually happens in industry, and the field has changed a lot in the last few years. In this episode we talk with David Holzm&#252;ller, a researcher at INRIA and one of the people behind TabArena, TabICL and RealMLP, about what the state of the art looks like right now and how to pick a model for your own data.</p><p>We cover the shift to TabPFN-style foundation models that learn to learn from whole tables, why TabArena was built and what earlier benchmarks got wrong, what Google's new TabFM means for the leaderboard, and when gradient boosted trees are still the right tool. David explains why LLMs struggle with tables, shares an early result comparing Claude Opus against TabICL on tiny datasets, and walks through how to embed text columns for tabular models. We also get into time series vs tabular data, the open research problems he thinks matter most, and why classical ML libraries are so bad out of the box.</p><div><hr></div><p><strong>Links</strong>:<br>TabArena:&nbsp;<a href="https://tabarena.ai">https://tabarena.ai</a></p><div><hr></div><p><strong>Topics</strong></p><ul><li><p>Tabular foundation models and in-context learning on tables</p></li><li><p>TabArena and Beyond Arena: building a benchmark that stays honest</p></li><li><p>TabFM, TabPFN, TabICL and the tradeoffs between them</p></li><li><p>When boosted trees and MLPs still win (large data, CPU, fast inference)</p></li><li><p>Why LLMs are inefficient on tabular data and where they might help</p></li><li><p>Embedding text columns with language models</p></li><li><p>Explainability, calibration and class imbalance</p></li><li><p>Time series vs tabular data</p></li><li><p>Open problems: invariances, synthetic data, uncertainty, scaling down</p></li><li><p>Where the field is heading in the next five years</p><div><hr></div></li></ul><p><strong>Chapters</strong></p><p>0:00 Intro<br>0:31 What changed in tabular ML: TabPFN-style foundation models<br>2:22 Which model to try first? TabArena and how it was built<br>5:14 What older benchmarks got wrong, and Beyond Arena<br>8:45 GPU AutoML vs foundation models<br>10:40 Reading the leaderboard: TabFM, TabPFN, TabICL and the tradeoffs<br>12:47 Calibration, class imbalance and small vs large data<br>19:45 Explainability for black-box tabular models<br>21:34 Why LLMs are bad at tabular data<br>25:39 Claude Opus 4.6 vs TabICL on tiny datasets<br>27:51 New classifiers, five-year outlook, real vs synthetic pretraining<br>33:13 Embedding text columns for tabular foundation models<br>36:13 Time series vs tabular data<br>39:59 When gradient boosted trees still win, and feature engineering<br>45:31 Open research problems and where the field is heading<br>52:54 Better MLPs and why classical defaults are bad out of the box</p><div><hr></div><ul><li><p><strong>Music</strong></p><ul><li><p>"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0</p></li></ul></li></ul>]]></content:encoded></item><item><title><![CDATA[Why You Can't Just Rent 1,000 GPUs | Charles Frye (Modal)]]></title><description><![CDATA[Charles Frye (Modal, ex-Weights & Biases, Berkeley PhD) joins Ravid and Allen to explain why modern AI research is bottlenecked by compute, and why simply buying more GPUs doesn't solve it.]]></description><link>https://www.the-information-bottleneck.com/p/why-you-cant-just-rent-1000-gpus-f85</link><guid isPermaLink="false">https://www.the-information-bottleneck.com/p/why-you-cant-just-rent-1000-gpus-f85</guid><dc:creator><![CDATA[Ravid Shwartz Ziv]]></dc:creator><pubDate>Tue, 01 Sep 2026 21:44:35 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/213775945/464fc9ee1997684ba30f32039241478e.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-aLzqkqgvz-E" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;aLzqkqgvz-E&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/aLzqkqgvz-E?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>Charles Frye (Modal, ex-Weights &amp; Biases, Berkeley PhD) joins Ravid and Allen to explain why modern AI research is bottlenecked by compute, and why simply buying more GPUs doesn't solve it. We cover the three problems every lab hits (underutilization, saturation, resource sharing), when companies should actually train their own models, why inference is a "bad algorithm" for today's hardware, NVIDIA's monopoly, the OpenAI/Hugging Face hack and what it says about open models, and whether we're in a compute bubble.</p><div><hr></div><p><strong>Key topics</strong></p><ul><li><p>AI infrastructure challenges and when to train your own models</p></li><li><p>GPU resource management and virtualization</p></li><li><p>Inference optimization and speculative decoding</p></li><li><p>The economics and future of AI hardware</p></li><li><p>Agents, sandboxing, and open-model security</p></li></ul><div><hr></div><p><strong>Chapters</strong><br>00:00 Intro<br>01:03 Why AI needs special-purpose compute<br>03:22 Buying vs renting GPUs: the three problems<br>07:15 Modal's approach, and doing more with less compute<br>09:46 Do we actually need to spend more? The conflict-of-interest question<br>13:08 Should companies train their own models?<br>14:47 Efficient fine-tuning and prompts as fast weights<br>17:37 Are we in a compute bubble?<br>20:21 Why inference will dominate compute (the SQLite analogy)<br>22:42 Speculative decoding<br>26:44 Why scaling inference is hard, and neuromorphic hardware<br>28:36 Why NVIDIA's monopoly persists<br>33:09 Inference chip startups and the hardware lottery<br>35:24 How Modal stays hardware-agnostic (GPU snapshot restore)<br>38:45 Will agentic coding erode CUDA's moat?<br>41:18 Running one agent vs thousands: sandboxing at scale<br>46:27 The OpenAI/Hugging Face hack and open models as defenders<br>52:28 Rogue AI, self-replication, and fast takeoff<br>56:09 What's next: evals, embodiment, edge inference<br>1:00:27 Modal is hiring (<a href="http://modal.jobs">modal.jobs</a>)</p><div><hr></div><ul><li><p><strong>Music</strong></p><ul><li><p>"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0</p></li></ul></li></ul>]]></content:encoded></item><item><title><![CDATA[Stella Biderman (EleutherAI) - Open Source, AI Safety, and Who We Can Trust]]></title><description><![CDATA[Stella Biderman, Executive Director of EleutherAI, joins us the week an OpenAI model autonomously broke out of its sandbox and hacked Hugging Face.]]></description><link>https://www.the-information-bottleneck.com/p/stella-biderman-eleutherai-open-source-d70</link><guid isPermaLink="false">https://www.the-information-bottleneck.com/p/stella-biderman-eleutherai-open-source-d70</guid><dc:creator><![CDATA[Ravid Shwartz Ziv]]></dc:creator><pubDate>Fri, 28 Aug 2026 03:44:08 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/213094873/14f2e043b714425cc6d392fad9e481f5.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-1c_wX4YdD2U" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;1c_wX4YdD2U&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/1c_wX4YdD2U?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>Stella Biderman, Executive Director of EleutherAI, joins us the week an OpenAI model autonomously broke out of its sandbox and hacked Hugging Face. Stella calls it what she thinks it is, an offensive cyber operation, and argues it&#8217;s part of a pattern: this is not the first containment failure at a frontier lab, and sandboxes have failed basically every time they&#8217;ve been tested for real.</p><p>So we spend a good chunk of the episode on what actual containment would look like. Stella&#8217;s argument is that the tools already exist, the labs just don&#8217;t use them: run dangerous capability evals on air-gapped networks with no route to the public internet, put the most sensitive testing in SCIF-style secure facilities, and treat model evaluation the way the security world treats classified systems rather than the way startups treat staging environments.</p><p>And yet Stella remains one of the world&#8217;s most prominent open-source advocates. From her perspective, the biggest risk isn&#8217;t the technology; it&#8217;s unchecked corporate power, and the only durable check on it is an independent scientific research establishment that doesn&#8217;t depend on the AI industry for its funding or its facts.</p><p>From there the conversation spans the geopolitics of Chinese open models and whether governments can restrict them, sovereign AI and what it would actually take for other countries to train their own models, why harnesses and UX drive more of AI&#8217;s perceived progress than raw intelligence, the AI-found counterexample to the Jacobian conjecture, and EleutherAI&#8217;s &#8220;Deep Ignorance&#8221; approach to making open-weight models safe by filtering hazardous knowledge out of pretraining.</p><div><hr></div><p><strong>key topics</strong></p><ul><li><p>AI governance and regulation</p></li><li><p>Cybersecurity incidents involving AI models</p></li><li><p>Open source AI safety and security</p></li><li><p>The role of independent research in AI safety</p></li><li><p>Legal and ethical considerations in AI development</p></li><li><div><hr></div><p><strong>Timeline</strong></p></li></ul><ul><li><p><strong>00:13</strong> &#8212; Intro: Stella Biderman and EleutherAI, a real non-profit in AI</p></li><li><p><strong>02:05</strong> &#8212; News of the week: Kimi K3, and OpenAI's model autonomously hacking Hugging Face</p></li><li><p><strong>05:49</strong> &#8212; "Frontier labs can't be trusted": repeated containment failures, air-gapped networks and SCIFs vs. sandboxes</p></li><li><p><strong>22:45</strong> &#8212; Can governments ban open or Chinese models? Import restrictions and the six-month open/closed gap</p></li><li><p><strong>27:05</strong> &#8212; Why Stella is still pro-open-source: unchecked corporate power as the real danger</p></li><li><p><strong>31:11</strong> &#8212; The opioid epidemic analogy: avoiding both regulatory failure and overcorrection</p></li><li><p><strong>34:57</strong> &#8212; Offense vs. defense: why open access to AI has empirically favored defenders</p></li><li><p><strong>37:28</strong> &#8212; Chinese labs, the CCP, and why safety and fine-tuning are low-prestige work in China</p></li><li><p><strong>42:19</strong> &#8212; Sovereign AI: does every country need its own foundation model?</p></li><li><p><strong>49:29</strong> &#8212; Sampling, harnesses, and why ChatGPT was really a UX breakthrough</p></li><li><p><strong>54:09</strong> &#8212; AI solves the Jacobian conjecture: domain data beats raw intelligence</p></li><li><p><strong>58:02</strong> &#8212; Safety is contextual, not a model property &#8212; and what HAL 9000 got right</p></li><li><p><strong>1:01:42</strong> &#8212; Is Stella optimistic about the future?</p></li><li><p><strong>1:02:50</strong> &#8212; Deep Ignorance, the science of AI training dynamics, and how to get involved with EleutherAI</p></li><li><div><hr></div></li></ul><ul><li><p><strong>Music</strong></p><ul><li><p>"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.</p></li></ul></li></ul>]]></content:encoded></item><item><title><![CDATA[Why AI Probably Won’t Cure Most Diseases in Ten Years]]></title><description><![CDATA[Since coding agents became much better at the end of last year, there has been a huge debate about what comes next.]]></description><link>https://www.the-information-bottleneck.com/p/why-ai-probably-wont-cure-most-diseases</link><guid isPermaLink="false">https://www.the-information-bottleneck.com/p/why-ai-probably-wont-cure-most-diseases</guid><dc:creator><![CDATA[Ravid Shwartz Ziv]]></dc:creator><pubDate>Wed, 26 Aug 2026 18:37:16 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!T8Ry!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feae27a0a-ff83-4c6e-9069-19ad87c9df54_1024x572.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!T8Ry!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feae27a0a-ff83-4c6e-9069-19ad87c9df54_1024x572.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!T8Ry!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feae27a0a-ff83-4c6e-9069-19ad87c9df54_1024x572.png 424w, https://substackcdn.com/image/fetch/$s_!T8Ry!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feae27a0a-ff83-4c6e-9069-19ad87c9df54_1024x572.png 848w, https://substackcdn.com/image/fetch/$s_!T8Ry!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feae27a0a-ff83-4c6e-9069-19ad87c9df54_1024x572.png 1272w, https://substackcdn.com/image/fetch/$s_!T8Ry!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feae27a0a-ff83-4c6e-9069-19ad87c9df54_1024x572.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!T8Ry!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feae27a0a-ff83-4c6e-9069-19ad87c9df54_1024x572.png" width="1024" height="572" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/eae27a0a-ff83-4c6e-9069-19ad87c9df54_1024x572.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:572,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:876287,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.the-information-bottleneck.com/i/212891536?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feae27a0a-ff83-4c6e-9069-19ad87c9df54_1024x572.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!T8Ry!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feae27a0a-ff83-4c6e-9069-19ad87c9df54_1024x572.png 424w, https://substackcdn.com/image/fetch/$s_!T8Ry!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feae27a0a-ff83-4c6e-9069-19ad87c9df54_1024x572.png 848w, https://substackcdn.com/image/fetch/$s_!T8Ry!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feae27a0a-ff83-4c6e-9069-19ad87c9df54_1024x572.png 1272w, https://substackcdn.com/image/fetch/$s_!T8Ry!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feae27a0a-ff83-4c6e-9069-19ad87c9df54_1024x572.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>Since coding agents became much better at the end of last year, there has been a huge debate about what comes next. If LLMs can transform coding so quickly, why not biology, drug discovery, or healthcare?</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.the-information-bottleneck.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Information Bottleneck is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Anthropic is now putting a lot of attention into these areas. Dario Amodei has gone much further, arguing that powerful AI could help cure most diseases within five to ten years.</p><p>I posted skeptically about this on social media, and I don&#8217;t remember ever receiving so many supportive DMs. Many came from people working in biology, running labs, or developing drugs.</p><p>I should say clearly that I&#8217;m not an expert in drug development. I have a bachelor&#8217;s degree in computational biology and a PhD in computational neuroscience, but I don&#8217;t currently work directly in biology. For this post, I spoke with several people who run biological labs or work at the intersection of AI and biology.</p><p>Of course, nobody really knows what AI systems will be capable of five years from now. But based on what we have today, I think there is a basic problem with the argument that progress in coding will transfer to biology.</p><h2>The verifier problem</h2><p>One useful way to think about AI tasks is to ask how easily an answer can be verified.</p><p>The fields where scaling and reinforcement learning have produced the most dramatic progress usually have cheap, fast feedback. Code has compilers and unit tests. Formal mathematics has proof checkers. Games have scores and clearly defined rules.</p><p>An AI can generate an answer, test it, learn from the result, and repeat this loop millions of times. This is extremely important. It allows the system to search for answers without a human having to check every attempt.</p><p>This is why I don&#8217;t find the argument &#8220;AI can solve hard mathematical problems, so drugs are next&#8221; very convincing. The mathematical result may be impressive, but it also shows what happens when you can generate a candidate and check it almost for free.</p><p>What happens when checking an answer takes three months? What if it costs hundreds of thousands of dollars? What if the experiment is noisy, and even after it ends you still don&#8217;t know whether the result will generalize to humans?</p><p>That is much closer to biology.</p><h2>A drug has to survive several different problems.</h2><p>A simplified drug-development process has at least four parts: mechanism, molecule, development, and market. A program must succeed at all four. Failure in any one of them can kill the entire project.</p><p><strong><span>Mechanism</span></strong> is the question of which biological process to intervene in.</p><p>We are worse at this than people think. High HDL cholesterol is associated with lower cardiovascular risk, yet several drugs that successfully raised HDL failed to improve clinical outcomes. The correlation was real. The intervention was wrong.</p><p>Vioxx worked as an anti-inflammatory, but it also increased cardiovascular risk. A drug can do exactly what it was designed to do and still cause serious harm somewhere else in the body. You cannot necessarily discover that by reading more papers or reasoning more carefully.</p><p><strong><span>Molecule</span></strong> is the question of what physical object can implement the intervention.</p><p>Liraglutide and semaglutide target the same receptor and are structurally similar, yet one has a half-life of roughly 13 hours while the other lasts about a week. The difference is not simply whether they bind to the receptor. Small molecular modifications change how they bind to albumin, resist degradation, and survive in circulation.</p><p>This is why success in binder design should not be confused with success in drug design. Binding matters, but a real drug must also reach the correct tissue, remain stable, avoid unwanted interactions, be manufacturable, and behave safely inside a complete organism.</p><p><strong><span>Development</span></strong> is where most of the time disappears.</p><p>Animal experiments take biological time. Some disease models require weeks or months before there is anything to measure. Then the experiment often has to be repeated to make sure the result was not noise.</p><p>Human trials are even harder. People must be recruited. Safety must be monitored. Some diseases progress slowly, so researchers need to wait long enough to see whether the treatment changes an actual outcome. According to the FDA, Phase II trials can take several months to two years, while Phase III trials often take one to four years.</p><p>You cannot solve all of this by buying more GPUs.</p><p><strong><span>The market</span></strong> also matters, even if people prefer not to discuss it. A known mechanism may be less risky but lead to a crowded market. A novel mechanism may have much greater value, but it also carries a much higher risk of failure. Some diseases have too few patients or too little purchasing power to attract enough investment.</p><p>There is no version of this process where everything becomes easy at the same time.</p><h2>Biology has verifiers, but they are bad ones.</h2><p>It would be too strong to say that biology has no verifiers. It has experiments.</p><p>The problem is that the important experiments are often slow, expensive, noisy, and incomplete.</p><p>Near the beginning of the pipeline, researchers can test whether a molecule binds to a target. But after that, the questions become much more difficult. Does it survive in the body? Does it reach the right tissue? Is it toxic? Does it work in an animal? Does the animal model actually predict what will happen in a person?</p><p>Eventually, you cannot test whether a compound is safe in humans without giving it to humans. You cannot identify rare side effects without studying enough people for long enough. You cannot determine whether a treatment delays a slowly progressing disease by running a faster simulation.</p><p>A clinical trial is not simply a very slow unit test. It is a noisy experiment conducted on a heterogeneous population, subject to ethical and operational constraints. In many cases, you get only a few serious attempts before the money or patent window runs out.</p><h2>What Anthropic actually showed</h2><p>None of this means that Anthropic&#8217;s recent result is unimportant.</p><p>Anthropic reported that Claude managed <em><span>de novo</span></em> protein-binder design campaigns and produced successful binders for 14 of 15 evaluated targets. External laboratories then produced and tested the designs. This is a genuinely impressive result. Anthropic also released the prompts, designs, and experimental data.</p><p>The company&#8217;s <a href="https://www.anthropic.com/research/Claude-accelerates-protein-design"><span>research post</span></a> is actually careful. It says that minibinders are not therapeutics and that a high-affinity binder is only the initial step toward a drug-like molecule.</p><p>But this caveat is not a small detail. It is the main point.</p><p>Binder design is one of the biological problems most compatible with current AI. You can generate many candidates computationally, rank them, and then test them with a relatively standardized experiment. It looks more like the problems where AI has already succeeded.</p><p>That is probably why we are seeing progress there first.</p><p>The mistake is assuming that because AI compressed this step, it will compress every later step at the same rate.</p><p>The scientific result is real. The questionable part is how the result changes as it moves upward: from designing protein binders to accelerating drug discovery to compressing decades of biological progress to curing most diseases within ten years.</p><p>At every step, a caveat disappears.</p><h2>The economics are different too.</h2><p>There is also a mismatch between the economics of frontier AI and those of drug development.</p><p>Software can be deployed globally, measured immediately, and updated continuously. Drug development requires long capital cycles and often results in many binary failures. A molecule can take years and cost hundreds of millions of dollars before a clinical trial shows it does not work.</p><p>This does not mean AI companies should avoid biology. But once a company begins to own and develop drugs, the economics start to look much more like pharma than software.</p><p>The promise to cure disease also does more than describe a research program. It provides a moral justification for enormous investments in compute, energy, and AI infrastructure. It makes the concentration of money and power around a few AI companies sound not only profitable yet necessary for humanity.</p><p>That doesn&#8217;t invalidate the scientific work. It does mean we should demand stronger evidence before accepting the larger story.</p><h2>What will happen?</h2><p>I use AI every day, and I am thrilled about projects that use it to search molecular space, design proteins, interpret experiments, and generate hypotheses.</p><p>AI will improve biological research. It will likely yield more drug candidates, and some may be much better than those we can design today.</p><p>But producing a prospective candidate starts the process. It does not finish it.</p><p>My view is that AI will substantially accelerate some parts of drug discovery without proportionally shortening the entire path from an idea to a safe, effective, widely available treatment.</p><p>The longest parts of that process are not simply periods when scientists sit around thinking too slowly. Researchers have to manufacture molecules, perturb living systems, recruit patients, observe outcomes, and wait for biology to reveal what happens.</p><p>This may change. Better simulations, robotic labs, stronger models, and more predictive animal or cellular systems could shorten some of these stages. AI may eventually become so accurate that it can select the rare molecule that satisfies efficacy, safety, delivery, stability, and manufacturing requirements on the first attempt.</p><p>But we have not seen evidence for that yet. And even doing it once would be very different from curing most diseases.</p><p>Anthropic&#8217;s binder result supports a narrower conclusion. It is less dramatic than curing most diseases, but still important:</p><p>AI is becoming extremely good at accelerating the parts of biology that resemble computation. Unfortunately, much of medicine is slow precisely because the decisive parts do not move quickly.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.the-information-bottleneck.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Information Bottleneck is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Why Deep Learning Finally Works on Tables | Frank Hutter (Prior Labs)]]></title><description><![CDATA[In this episode, Frank Hutter joins us to talk about TabPFN and why tabular data is suddenly the hottest problem in deep learning.]]></description><link>https://www.the-information-bottleneck.com/p/why-deep-learning-finally-works-on-6ab</link><guid isPermaLink="false">https://www.the-information-bottleneck.com/p/why-deep-learning-finally-works-on-6ab</guid><dc:creator><![CDATA[Ravid Shwartz Ziv]]></dc:creator><pubDate>Mon, 24 Aug 2026 04:18:33 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/212499375/485e6d4eac7eaee7e7b10047912b3a23.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-LHuIt8U0GZw" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;LHuIt8U0GZw&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/LHuIt8U0GZw?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>In this episode, Frank Hutter joins us to talk about TabPFN and why tabular data is suddenly the hottest problem in deep learning. Frank is a professor at the University of Freiburg and spent 15 years building the AutoML field before founding Prior Labs, which SAP just acquired for over a billion dollars.</p><p>We get into why deep learning failed on tables for a decade and what in-context learning changed, how TabPFN is trained entirely on synthetic data, and why a model that never saw a real time series ended up beating specialized forecasting models. Frank also explains the architecture tricks behind scaling from 10,000 to a million rows, where LLMs fit into data science (and where they embarrassingly don't), and what happens to XGBoost from here.</p><p>Beyond the research, Frank talks about the jump from professor to co-CEO, why he refused to merge his 45-person team into SAP's 110,000 employees, the open-weights licensing debate, and the case for building a frontier lab in Freiburg rather than San Francisco.</p><div><hr></div><p><strong>key topics</strong></p><ul><li><p>The role of foundation models in tabular data</p></li><li><p>Impact of SAP acquisition on Pro Labs</p></li><li><p>The evolution of AutoML and hyperparameter optimization</p></li><li><p>Challenges and solutions for large context in models</p></li><li><p>Open source models and licensing strategies</p></li><li><p>The importance of independence for startup agility</p></li><li><p>Future directions in AI for science and medicine</p></li><li><div><hr></div><p>00:00 Intro<br>00:34 The SAP acquisition and staying independent<br>07:39 Why tabular data is the next big thing in deep learning<br>14:19 What makes tabular data hard<br>19:14 AutoML, AutoGluon, and fifteen years of hyperparameter tuning<br>28:27 Scaling TabPFN: context limits and architectures<br>34:35 Agentic data science and LLMs<br>39:30 Online learning, time series, and Bayesian inference in a forward pass<br>47:05 Open weights and the license debate<br>54:51 Will LLMs and tabular models merge?<br>1:00:01 From academia to startup<br>1:09:42 Why build in Europe<br>1:12:53 Audience questions and hiring</p></li></ul><div><hr></div><ul><li><p><strong>Music</strong></p><ul><li><p>"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.</p></li></ul></li></ul>]]></content:encoded></item><item><title><![CDATA[Surya Ganguli: The Physics of Intelligence]]></title><description><![CDATA[Surya Ganguli is a professor at Stanford and VP at General Catalyst, working at the intersection of physics, neuroscience, and AI.]]></description><link>https://www.the-information-bottleneck.com/p/surya-ganguli-the-physics-of-intelligence-051</link><guid isPermaLink="false">https://www.the-information-bottleneck.com/p/surya-ganguli-the-physics-of-intelligence-051</guid><dc:creator><![CDATA[Ravid Shwartz Ziv]]></dc:creator><pubDate>Mon, 17 Aug 2026 21:23:17 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/211626291/e6ca9ceafefc48f1e4b16b3a23addc36.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-XsQxIYzK-OM" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;XsQxIYzK-OM&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/XsQxIYzK-OM?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>Surya Ganguli is a professor at Stanford and VP at General Catalyst, working at the intersection of physics, neuroscience, and AI. He started in string theory, moved to theoretical neuroscience, and now uses tools from statistical physics to understand both brains and neural networks.</p><p>We talk about why deep learning theory is finally catching up to practice, &nbsp;including his group's recent work explaining neural scaling laws, and why smarter data selection could beat them entirely. He also tells the origin story of diffusion models, which were invented in his lab as an attempt to violate the second law of thermodynamics.</p><p>The second half turns to the brain: what happens to a mouse's sense of self on ketamine, how stimulating a handful of neurons can induce hallucinations, and a method his lab developed to get a neuron deep in a monkey's brain to describe, in English, what makes it fire.</p><p>We close on where he thinks AI is going wrong: models train on ten trillion tokens while humans hear a hundred million words, because we don't teach children with gradients; we tell them the algorithm.</p><div><hr></div><p><strong><br>key topics</strong></p><ul><li><p>Connections between physics, neuroscience, and AI</p></li><li><p>Emergent properties in complex systems</p></li><li><p>Scaling laws in language models</p></li><li><p>Data efficiency and pruning in AI</p></li><li><p>Neuroscience insights into consciousness and self</p></li><li><p>The future of AI and brain modeling</p></li><li><div><hr></div><p><strong>Chapters</strong></p><p><strong>00:00 </strong>Introduction to Surya Ganguli</p><p><strong>00:57 </strong>Surya's Background: From String Theory to Neuroscience</p><p><strong>02:22 </strong>Emergent Properties in Physics, Neuroscience, and AI</p><p><strong>03:16 </strong>Energy Landscapes and Loss Landscapes in High Dimensions</p><p><strong>04:07 </strong>Why Local Minima Don't Exist in High-Dimensional AI</p><p><strong>05:22 </strong>Gradient-Based vs. Gradient-Free Learning Methods</p><p><strong>08:21 </strong>AI in Mathematics and Drug Discovery: Opportunities and Challenges</p><p><strong>13:48 </strong>Scaling Laws and Data Efficiency in Language Models</p><p><strong>18:10 </strong>Properties of Data that Affect Scaling Laws</p><p><strong>22:04 </strong>Constructing Non-Redundant Data Sets for Better Learning</p><p><strong>24:32 </strong>Theory vs. Empirical Results in AI Research</p><p><strong>32:19 </strong>Fundamental Components of Deep Learning: Are They Changing?</p><p><strong>34:31 </strong>Future Paradigms in AI Beyond Current Models</p><p><strong>37:22 </strong>Teaching AI and Humans: Paradigm Shifts in Learning</p><p><strong>41:37 </strong>Consciousness, Self, and the Brain: Surya's Perspectives</p><p><strong>49:49 </strong>Neuroscience and AI: Understanding the Brain and Consciousness</p><p><strong>01:02:03 </strong>Understanding the Brain: Challenges and Opportunities</p><p><strong>01:09:21 </strong>Brain-Computer Interfaces and AI in Neuroscience</p></li><li><div><hr></div><p><strong>Music</strong></p><ul><li><p>"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.</p></li></ul></li></ul>]]></content:encoded></item><item><title><![CDATA[Text Diffusion Models with Brendan O'Donoghue (Google DeepMind)]]></title><description><![CDATA[Brendan O&#8217;Donoghue, research director at Google DeepMind, makes the case for text diffusion as a real alternative to autoregressive generation.]]></description><link>https://www.the-information-bottleneck.com/p/text-diffusion-models-with-brendan</link><guid isPermaLink="false">https://www.the-information-bottleneck.com/p/text-diffusion-models-with-brendan</guid><dc:creator><![CDATA[Ravid Shwartz Ziv]]></dc:creator><pubDate>Fri, 14 Aug 2026 15:28:02 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/211191730/62f8f0e6debb20e7b3cd1bd09b2c2ed7.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-o9_g86_dlLg" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;o9_g86_dlLg&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/o9_g86_dlLg?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>Brendan O&#8217;Donoghue, research director at Google DeepMind, makes the case for text diffusion as a real alternative to autoregressive generation. He walks through how discrete diffusion works, why diffusion samples are far more diverse and what that unlocks for RL, where the Gemma diffusion model actually stands against frontier models, and why the whole training and serving stack being hyper-optimized for autoregression is the main thing holding the approach back. The conversation also covers hardware trends favoring flops over bandwidth, AGI timelines and real-world bottlenecks, and why he thinks RL is still underhyped.</p><div><hr></div><p><strong>Key topics</strong><br>- Discrete diffusion for text vs autoregressive generation<br>- Why diffusion samples are more diverse, and what that unlocks for RL<br>- Where diffusion already wins: latency, on-device, robotics<br>- Why serving cost, not quality, is the real blocker<br>- RL as the most underhyped area in AI</p><div><hr></div><p><strong>Timeline</strong><br>00:00 Introduction<br>00:50 What diffusion models are and how text diffusion works<br>04:40 Why Brendan bet on text diffusion in 2023<br>07:15 Diversity, creativity, and why it helps RL<br>11:00 The best diffusion LLM today and the gap to frontier models<br>14:25 Latency, serving cost, and why it needs more chips<br>17:14 Where diffusion already wins: on-device, robotics, battery<br>20:14 One model, two modes: diffusion for thinking, AR for answering<br>22:24 Samplers and the stuttering problem<br>26:27 Theory, BERT, and why now is a good time to work on this<br>31:48 Pipelines built for autoregression, and continuous diffusion<br>35:35 Hardware: flops vs bandwidth<br>39:49 AGI timelines and real-world bottlenecks<br>50:15 Is AI engineering or science?<br>54:14 Most overhyped and most underhyped ideas<br>58:35 RL on diffusion, value functions, and exploration<br>1:07:30 Go download the model and break it</p><div><hr></div><p><strong>Music</strong></p><ul><li><p>&#8220;Kid Kodi&#8221; - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.</p></li></ul>]]></content:encoded></item><item><title><![CDATA[Nathan Lambert: Inside Post-Training and the Open Model Fight]]></title><description><![CDATA[Nathan Lambert spent three years as post-training lead at Ai2, where he built the OLMo models, and he writes Interconnects, one of the most-read technical newsletters in AI.]]></description><link>https://www.the-information-bottleneck.com/p/nathan-lambert-inside-post-training</link><guid isPermaLink="false">https://www.the-information-bottleneck.com/p/nathan-lambert-inside-post-training</guid><dc:creator><![CDATA[Ravid Shwartz Ziv]]></dc:creator><pubDate>Sat, 08 Aug 2026 17:11:36 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/210368358/fa5b70e3c3fb9fb022f485b96b8c3db5.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-G3zanJBcPuo" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;G3zanJBcPuo&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/G3zanJBcPuo?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>Nathan Lambert spent three years as post-training lead at Ai2, where he built the OLMo models, and he writes Interconnects, one of the most-read technical newsletters in AI. He left Ai2 in June and is now working on a new project. He&#8217;s also the author of the RLHF book. We talked a lot about open models, their capabilities, and why they are better than he expected. We get into what that means over the next two to five years, why he thinks recursive self-improvement is overblown, what the market for training environments actually looks like now, and why he expects Anthropic&#8217;s famously open internal culture to break after its IPO.</p><div><hr></div><p><strong>Key Topics</strong></p><ul><li><p>Open vs closed models and who actually captures the value</p></li><li><p>Anthropic and OpenAI as opposite cultures, and the talent concentration problem</p></li><li><p>Boom vs bubble, and why token spend hasn&#8217;t produced 10x better products</p></li><li><p>Continual learning, RSI skepticism, and what Nathan wants to work on next</p></li><li><p>What the open ecosystem needs economically to survive</p></li></ul><div><hr></div><p><strong>Timeline</strong></p><p><strong>00:00</strong> Intro<br><strong>00:27</strong> Open vs closed models, and who actually captures the value<br><strong>05:12</strong> China, harnesses, and where the real training leverage sits<br><strong>08:40</strong> Sovereign compute and the national security case for building models<br><strong>11:18</strong> Uncensored open weights and the bioweapon question<br><strong>14:29</strong> Anthropic vs OpenAI, ideology and politics<br><strong>19:35</strong> The Mythos ban and the Fable 5 delays<br><strong>24:30</strong> The AGI narrative, the talent drain, and antitrust<br><strong>28:12</strong> Why researchers join Anthropic, and the open Slack culture<br><strong>34:04</strong> Nathan&#8217;s next 12 months: character training and big RL runs<br><strong>37:55</strong> Continual learning, RSI, and why Nathan is skeptical<br><strong>43:19</strong> Boom or bubble, tokens vs GPUs<br><strong>45:12</strong> Why all that token spend never produced 10x products<br><strong>48:38</strong> Job displacement and the small-business future<br><strong>52:49</strong> Robotics, world models, and why multimodal lags<br><strong>57:44</strong> What the open ecosystem should actually do<br><strong>1:03:17</strong> Why NVIDIA isn&#8217;t building a frontier model<br><strong>1:07:34</strong> The RLHF book, and whether RLHF still matters<br><strong>1:11:06</strong> GRPO vs PPO and on-policy distillation</p><div><hr></div><p><strong>Music</strong></p><ul><li><p>&#8220;Kid Kodi&#8221; - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.</p></li></ul>]]></content:encoded></item><item><title><![CDATA[Daphne Koller - The Future of AI in Biology and Drug Discovery]]></title><description><![CDATA[Daphne Koller wrote the book that many of us learned probabilistic graphical models from, founded Coursera, and now runs insitro, which is trying to make drug discovery a machine-learning problem.]]></description><link>https://www.the-information-bottleneck.com/p/daphne-koller-the-future-of-ai-in</link><guid isPermaLink="false">https://www.the-information-bottleneck.com/p/daphne-koller-the-future-of-ai-in</guid><dc:creator><![CDATA[Ravid Shwartz Ziv]]></dc:creator><pubDate>Tue, 04 Aug 2026 04:12:07 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/209734077/cb95ff1ad7811e3e3f914f934c20f485.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-EACwotMLWog" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;EACwotMLWog&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/EACwotMLWog?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>Daphne Koller wrote the book that many of us learned probabilistic graphical models from, founded Coursera, and now runs insitro, which is trying to make drug discovery a machine-learning problem.</p><p>We start with the bitter lesson. She agrees with most of it and then says where it stops working: biology doesn&#8217;t have enough data, structure is how people understand anything, and making a drug is a question about an intervention that hasn&#8217;t happened yet, not a pattern in data you already have.</p><p>Most of the episode is about why drug discovery is hard. Ninety percent of drugs that reach the clinic fail, and mostly not because the molecule was bad. The molecule usually does what it was designed to do. It just turns out the thing it was designed to do had nothing to do with the disease. Only 22% of diseases have any approved drug at all, and she calls that an upper bound on what we understand, not a lower bound.</p><p>She also gets into what agents are and aren&#8217;t good for in a wet lab, why cells don&#8217;t grow faster no matter how many GPUs you point at them, what it would take to have real foundation models for biology, and why almost all of biology is still out of distribution.</p><p>Plus GLP-1s and what human data keeps teaching us, whether AI can make the kind of leap that turned a bacterial immune system into CRISPR, and what she&#8217;d build if she were starting Coursera today.</p><div><hr></div><p><strong>Key Topics</strong></p><ul><li><p>The impact of scaling and data in machine learning</p></li><li><p>The importance of structure and causality in AI</p></li><li><p>Challenges in drug discovery and biological understanding</p></li><li><p>The role of foundation models in biology</p></li><li><p>Ethical considerations in AI and biomedical research</p></li></ul><div><hr></div><p><strong>Chapters</strong></p><p>00:00 Introduction to Machine Learning and Drug Discovery</p><p>02:00 The Bitter Lesson and Its Implications</p><p>06:48 Challenges in Drug Design and Discovery</p><p>11:48 Ethical Considerations in Human Research</p><p>17:20 The Drug Discovery Pipeline Explained</p><p>29:30 Integrating AI in Experimental Design</p><p>35:38 The Role of Human Judgment in Drug Design</p><p>37:14 Future of Drug Design: Efficiency vs. Automation</p><p>39:37 Challenges in AI and Data Availability for Biology</p><p>41:08 Foundation Models: Potential and Limitations</p><p>43:39 Causality in Biological Data: Importance and Challenges</p><p>45:18 Creativity vs. Understanding in Drug Design</p><p>48:17 Balancing Investments in Data, Algorithms, and Experiments</p><p>50:07 The Value of Simulations in Drug Discovery</p><p>52:03 Mathematical Frameworks in Biology: Utility and Limitations</p><p>54:14 The Future of Drug Discovery: Optimism and Innovations</p><p>56:28 The Impact of Coursera on Education</p><p>01:00:33 The Role of Universities in Lifelong Learning</p><p>01:04:06 Connecting Dots: The Fun of Variety in Work</p><p>01:05:46 Optimism for the Future of Drug Discovery</p><div><hr></div><p><strong>Music</strong></p><ul><li><p>&#8220;Kid Kodi&#8221; - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.</p></li></ul>]]></content:encoded></item><item><title><![CDATA[RL Was Broken at Every Level - With Joseph Suarez (PufferAI)]]></title><description><![CDATA[Joseph Suarez on why deep RL stalled, and what fixing the code actually bought]]></description><link>https://www.the-information-bottleneck.com/p/rl-was-broken-at-every-level-with</link><guid isPermaLink="false">https://www.the-information-bottleneck.com/p/rl-was-broken-at-every-level-with</guid><dc:creator><![CDATA[Ravid Shwartz Ziv]]></dc:creator><pubDate>Thu, 30 Jul 2026 19:37:10 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/209156830/1cabf24a91d6029f250f83b9ad66eb3b.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-8Sv4QVbOAWA" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;8Sv4QVbOAWA&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/8Sv4QVbOAWA?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>In this episode, Joseph Suarez from PufferAI explains why he thinks RL never had an algorithm problem, but it had a code problem. Every part of the standard RL stack was running about a thousand times slower than it should have been, and once that got fixed, problems that used to take months started getting solved in seconds on one GPU. We talk about what makes a simulator good for RL, why most of their sims run on CPU, what he wants to do with scientific simulation, and why he open sources all of it instead of writing papers.</p><div><hr></div><p><strong>Key topics</strong></p><ul><li><p>Types of RL and their applications</p></li><li><p>Challenges in scaling reinforcement learning</p></li><li><p>The role of simulators and hardware in RL</p></li><li><p>RL in gaming: from chess to complex games like NetHack and RuneScape</p></li><li><p>Future directions: scientific simulation and biological modeling</p></li></ul><div><hr></div><p><strong>Chapters</strong></p><p><strong>00:00 - </strong>Introduction to RL and Puff AI</p><p><strong>01:50 - </strong>Different settings for RL: Games, Robots, Finance</p><p><strong>04:10 - </strong>RL in LM and other domains</p><p><strong>07:00 - </strong>Challenges and solutions in RL scaling</p><p><strong>09:55 - </strong>Building fast, efficient simulators</p><p><strong>15:10 - </strong>RL for scientific research and simulation</p><p><strong>19:57 - </strong>RL in complex games: NetHack, RuneScape, Dwarf Fortress</p><p><strong>29:55 - </strong>Future of RL: Scientific discovery and beyond</p><div><hr></div><p><strong>Resources</strong></p><p>Puff AI - Official Site -  <a href="https://puffer.ai">https://puffer.ai</a></p><p>NetHack -  <a href="https://www.nethack.org/">https://www.nethack.org/</a></p><p>RuneScape -  <a href="https://www.runescape.com/">https://www.runescape.com/</a></p><p>Dwarf Fortress - <a href="http://www.bay12games.com/dwarves/">http://www.bay12games.com/dwarves/</a></p><p>OpenAI Gym - <a href="https://github.com/openai/gym">https://github.com/openai/gym</a></p><div><hr></div><p><strong>Music</strong></p><ul><li><p>&#8220;Kid Kodi&#8221; - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.</p></li></ul>]]></content:encoded></item><item><title><![CDATA[The Model Found a Way Out  -  with Florian Brand (Prime Intellect)  ]]></title><description><![CDATA[Florian Brand builds evals at Prime Intellect.]]></description><link>https://www.the-information-bottleneck.com/p/the-model-cheated-with-florian-brand</link><guid isPermaLink="false">https://www.the-information-bottleneck.com/p/the-model-cheated-with-florian-brand</guid><dc:creator><![CDATA[Ravid Shwartz Ziv]]></dc:creator><pubDate>Mon, 27 Jul 2026 14:43:41 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/208695274/4e22fdab66c93770074dcb7b641e34f5.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-lrfMxsDGeW4" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;lrfMxsDGeW4&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/lrfMxsDGeW4?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>Florian Brand builds evals at Prime Intellect. The premise of the conversation is that writing a benchmark is the easy part now. Keeping the model from cheating it is the job, and it takes longer than the benchmark itself.</p><p>We get into why he thinks you can&#8217;t evaluate a model apart from the CLI it runs in, what happens to statistics when a single run costs five figures, and whether the feeling that a model just works can ever become a number.</p><p>He also has a few stories about agents finding their way around the scoring that are worth hearing cold.</p><div><hr></div><p><strong>Timeline</strong></p><ul><li><p>00:13 Intro</p></li><li><p>01:00 What evals are for</p></li><li><p>04:05 Agentic benchmarks</p></li><li><p>07:10 Kimi K2 and model diversity</p></li><li><p>08:23 Long-horizon coding tasks</p></li><li><p>10:29 Building a benchmark</p></li><li><p>12:15 MirrorCode</p></li><li><p>14:27 Rubrics and LLM judges</p></li><li><p>16:30 The cost of expert labelers</p></li><li><p>17:49 Long runs and variance</p></li><li><p>19:44 Evaluating the harness</p></li><li><p>24:29 Chinese labs building CLIs</p></li><li><p>30:00 More reward hacking</p></li><li><p>37:45 Tau-bench and economic tasks</p></li><li><p>39:43 Benchmaxxing and GLM 5.2</p></li><li><p>45:15 Statistics and cost</p></li><li><p>47:56 Frontier convergence</p></li><li><p>52:04 Misuse in open and closed models</p></li><li><p>55:35 Self-improvement</p></li><li></li></ul><div><hr></div><p><strong>Music</strong></p><ul><li><p>&#8220;Kid Kodi&#8221; - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.</p></li></ul><p><strong>About</strong></p><p>The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.</p>]]></content:encoded></item><item><title><![CDATA[Pierre-Carl Langlais on Building Models from Data You Can Account For]]></title><description><![CDATA[Most labs build language models by scraping the web and filtering afterward.]]></description><link>https://www.the-information-bottleneck.com/p/pierre-carl-langlais-on-building</link><guid isPermaLink="false">https://www.the-information-bottleneck.com/p/pierre-carl-langlais-on-building</guid><dc:creator><![CDATA[Ravid Shwartz Ziv]]></dc:creator><pubDate>Thu, 23 Jul 2026 13:56:18 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/208193080/b28d119102152869c01bf6c865df7639.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-vgoO320MNT8" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;vgoO320MNT8&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/vgoO320MNT8?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>Most labs build language models by scraping the web and filtering afterward. Pierre-Carl Langlais runs it the other way around. At Pleias, the French-German lab he co-founded, the models are built from data he can actually account for, which in practice means open and public-domain sources plus a lot of synthetic data the lab generates itself. It sounds like a self-imposed handicap. It mostly isn&#8217;t. One of their models is a 600 million parameter system that runs live inside the Paris subway&#8217;s monitoring pipeline.</p><p>We cover the SYNTH pretraining dataset and why he thinks &#8220;ethical data&#8221; has to mean more than copyright-free. He explains why barely 2% of their Common Corpus appears in typical web crawls, and why that gap is really a preservation problem. From there, he gets blunt about benchmark maxing and whether GLM really earns its Opus-class reputation. He also argues that the quiet move by closed labs to hide reasoning traces is mostly about claiming ownership of model outputs. He&#8217;s skeptical of sovereign AI, and not shy about how Mistral drifted from frontier research toward French corporate consulting. We finish on NVIDIA&#8217;s persona datasets and the odd idea of training on the conditions that produced a text rather than the text itself.</p><h3>Timeline</h3><ul><li><p>(00:02) Welcome and introductions</p></li><li><p>(00:49) Why synthetic data matters, and the SYNTH set</p></li><li><p>(04:15) Three reasons to control your training data</p></li><li><p>(07:18) What &#8220;ethical data&#8221; actually means</p></li><li><p>(11:08) How Common Corpus got built, from Wikipedia to PDFs</p></li><li><p>(16:35) Agentic harnesses and synthetic data</p></li><li><p>(20:03) Evaluating data when you train on reasoning traces</p></li><li><p>(25:27) General versus specialized pretraining</p></li><li><p>(27:08) Benchmark maxing and the GLM question</p></li><li><p>(31:51) Getting diversity in, and the NVIDIA personas</p></li><li><p>(35:02) Hidden reasoning traces and the fight over model IP</p></li><li><p>(38:17) Mid-training and the &#8220;It&#8217;s All Training&#8221; thesis</p></li><li><p>(41:47) Can small models actually compete</p></li><li><p>(45:01) Cybersecurity and Europe&#8217;s strategic gap</p></li><li><p>(47:08) Do you need a big model to orchestrate the small ones</p></li><li><p>(52:08) Sovereign AI and the limits of national champions</p></li><li><p>(56:42) Scaling laws when you control the data</p></li><li><p>(01:00:41) The NVIDIA persona datasets</p></li><li><p>(01:04:52) What you actually do with synthetic personas</p></li><li><p>(01:08:22) Closing thoughts</p></li></ul><p><strong>Music</strong></p><ul><li><p>&#8220;Kid Kodi&#8221; - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.</p></li></ul><p><strong>About</strong></p><p>The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.</p>]]></content:encoded></item><item><title><![CDATA[Dhruv Batra: The Browser Is a Robotics Problem - From Embodied AI at Meta to Web Agents at Yutori]]></title><description><![CDATA[Dhruv Batra spent years leading Embodied AI at Meta, training virtual robots to navigate photorealistic 3D scans of real buildings with pure reinforcement learning.]]></description><link>https://www.the-information-bottleneck.com/p/dhruv-batra-the-browser-is-a-robotics</link><guid isPermaLink="false">https://www.the-information-bottleneck.com/p/dhruv-batra-the-browser-is-a-robotics</guid><dc:creator><![CDATA[Ravid Shwartz Ziv]]></dc:creator><pubDate>Mon, 20 Jul 2026 15:41:28 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/207788591/6419f0ed61bacb531216168bdd697b21.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p></p><div id="youtube2-TAfcqUydqM0" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;TAfcqUydqM0&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/TAfcqUydqM0?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>Dhruv Batra spent years leading Embodied AI at Meta,  training virtual robots to navigate photorealistic 3D scans of real buildings with pure reinforcement learning. Then he left to co-found Yutori and build agents for a very different environment: the web browser.</p><p>In this episode, Dhruv explains why he sees these as the same problem. Web agents, in his framing, are robots that act in a browser (pixels in, actions out), and the web turns out to be just as messy an environment as the physical world.</p><p>Along the way, we cover his definition of intelligence as &#8220;navigation in idea space,&#8221; why robotics is lagging LLMs, the sim-to-real gap and why you can&#8217;t fake friction coefficients, the teleoperation counterexample to the &#8220;it&#8217;s a sensor problem&#8221; argument, and his provocative claim that under the current paradigm, we solved machine learning and didn&#8217;t even realize it. He also makes the case for why the scaling hypothesis isn&#8217;t falsifiable, why JEPA-style arguments deserve to be grappled with, how Yutori trains its Navigator models with RL on <em>live</em> websites, and what happens to the ad-supported web when agents, not eyeballs, do the browsing.</p><h2>Timeline</h2><p><strong>00:01</strong> &#8212; Intro<br><strong>00:54</strong> &#8212; What embodied AI actually means<br><strong>06:47</strong> &#8212; Intelligence as navigation in idea space<br><strong>13:26</strong> &#8212; Habitat: training robots with pure RL, no maps<br><strong>20:04</strong> &#8212; Why robotics is behind LLMs<br><strong>28:24</strong> &#8212; Sim-to-real: what you can and can&#8217;t fake<br><strong>33:34</strong> &#8212; &#8220;We solved ML and nobody noticed&#8221;<br><strong>37:12</strong> &#8212; Leaving Meta, founding Yutori<br><strong>43:21</strong> &#8212; Web agents: screenshots in, actions out<br><strong>48:15</strong> &#8212; Why the web won&#8217;t rebuild itself for agents<br><strong>53:32</strong> &#8212; Training Navigator: RL on live websites<br><strong>1:01:04</strong> &#8212; Who pays for the web when agents browse?<br><strong>1:09:17</strong> &#8212; What Yutori means, closing thoughts</p><p><strong>Music</strong></p><ul><li><p>&#8220;Kid Kodi&#8221; - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.</p></li></ul><p><strong>About</strong></p><p>The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.</p>]]></content:encoded></item><item><title><![CDATA[How to Turn Research Into Billion-Dollar Companies, with Ion Stoica]]></title><description><![CDATA[Ion Stoica has done what almost no academic ever does &#8212; repeatedly turned university research into billion-dollar companies.]]></description><link>https://www.the-information-bottleneck.com/p/how-to-turn-research-into-billion</link><guid isPermaLink="false">https://www.the-information-bottleneck.com/p/how-to-turn-research-into-billion</guid><dc:creator><![CDATA[Ravid Shwartz Ziv]]></dc:creator><pubDate>Thu, 16 Jul 2026 19:55:47 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/207329369/095ba6b39cbe1ec09976bf810097fa09.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-QkEYr5jW4BE" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;QkEYr5jW4BE&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/QkEYr5jW4BE?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>Ion Stoica has done what almost no academic ever does &#8212; repeatedly turned university research into billion-dollar companies. He co-founded Databricks (now valued at over $100 billion), Anyscale, Arena AI and Conviva, while his Berkeley lab produced the open source projects the entire AI industry runs on: Ray, vLLM, and SGLang.</p><p>In this episode, we ask him how it&#8217;s actually done. His answer is surprisingly unromantic: solve a problem people already care about, build an artifact good enough that they adopt it, and pay attention to the moment users start asking &#8220;who maintains this after the students graduate?&#8221; - that&#8217;s when a project becomes a company. He&#8217;s also insistent that the credit belongs to his students.</p><p>From there, the conversation goes deep into what he&#8217;s watching now: why the AI stack has become an order of magnitude more complex than the Hadoop/Spark era, why maximizing GPU utilization is &#8220;the name of the game&#8221; for any enterprise, and why coding agents will struggle with distributed systems long after they&#8217;ve mastered web apps. He shares a memorable reward-hacking story &#8212; a load balancer that maximized throughput by dropping requests &#8212; explains why the gap between open and closed models sits at about six months, and closes with his case for regulating AI by outcomes, not capabilities.</p><p><strong>Timeline</strong></p><ul><li><p>00:00 &#8212; Introduction: welcoming Ion Stoica</p></li><li><p>01:21 &#8212; The playbook: how research projects become companies</p></li><li><p>05:22 &#8212; Will vLLM and SGLang stay open source?</p></li><li><p>07:47 &#8212; The real bottleneck in the AI stack: complexity, not just hardware</p></li><li><p>14:31 &#8212; Should algorithms follow infrastructure, or the other way around?</p></li><li><p>16:13 &#8212; Can AI coding tools write distributed systems and GPU kernels?</p></li><li><p>21:09 &#8212; Verifiers, harnesses, and the limits of outsourcing understanding</p></li><li><p>25:41 &#8212; Reward hacking: the load balancer that dropped requests</p></li><li><p>25:58 &#8212; How should enterprises consume GPUs? Utilization as the name of the game</p></li><li><p>30:23 &#8212; GPU scarcity: will the compute crunch ever end?</p></li><li><p>35:27 &#8212; Hyper-optimization and the risk of locking in today&#8217;s architectures</p></li><li><p>37:17 &#8212; Open vs. closed models: why every company wants to own the stack</p></li><li><p>40:35 &#8212; The six-month gap, and the rising cost of training frontier models</p></li><li><p>43:58 &#8212; Kimi, Qwen, and who&#8217;s incentivized to keep open models alive</p></li><li><p>45:39 &#8212; Regulation: outcomes, not capabilities</p></li><li><p>47:41 &#8212; Self-regulation, concentration of power, and auditing open models</p></li><li><p>48:32 &#8212; Wrap-up</p></li></ul><p><strong>Music</strong></p><ul><li><p>&#8220;Kid Kodi&#8221; - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.</p></li></ul><p><strong>About</strong></p><p>The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.</p>]]></content:encoded></item></channel></rss>