The Information Bottleneck
The Information Bottleneck
Continual Learning Is the Next Bottleneck | Rohan Anil (Core Automation )
0:00
-1:11:09

Continual Learning Is the Next Bottleneck | Rohan Anil (Core Automation )

Rohan Anil spent eleven and a half years at Google, where he went from writing memory allocators to large-scale linear solvers, then optimization at Google Brain, where he co-developed distributed Shampoo and led optimization for PaLM and Gemini pre-training, including the work that produced Gemini Flash. He then joined Anthropic's pre-training team, and left before the IPO to co-found Core Automation with Jerry Tworek (ex-VP of Research at OpenAI). We talk with him about how Brain worked at its peak, why he left two of the world's best labs, and what he thinks is missing from today's models.

Rohan's view is that pre-training and RL were split by organizational convenience rather than by science. Pre-training builds a prior, and RL sharpens it to the tasks we care about, and neither gives a model a way to absorb new data or learn from its own experience once it is deployed. Post-training more every day plateaus, on-policy distillation plateaus, and in-context learning only goes as far as the context does. He argues the next architecture needs better ways to fold in new knowledge at inference time, and that this is a fundamental optimization question rather than a harness-engineering one.

We also get into why coding agents still fail on low-level systems work, his take on Muon, why second-order methods matter once you leave the noise-dominated regime, and why nobody can yet use a few million GPUs for a single training run.


Timeline

00:00 Intro
01:09 From computer vision to Google systems engineering
02:37 Large-scale linear solvers and sparse features
06:01 Getting into optimization: SDCA and Yonghui Wu's team
07:29 Joining the Shampoo crew
09:35 The Google Brain ethos, and why 2017 to 2019 was special
14:23 Is open research going to keep winning?
15:53 Frontier models are only as good as the prior you give them
17:30 Missing the language model wave, then Common Crawl and online distillation
18:31 Paternity leave, DALL-E Mini, and the 14 days that became two years
20:32 PaLM, Gemini pre-training, and Gemini Flash
23:59 The Shampoo origin story: Tomer Koren's two-week proof
26:54 Why leave Google for Anthropic
30:20 Why leave Anthropic for a startup
31:33 Meeting Jerry Tworek at Dolores Park
34:00 What Core Automation is building
36:17 Continual learning and the pre-training vs RL split
39:11 Why coding agents fail at kernels and low-level pipelines
42:00 The QR factorization kernel competition and reward hacking
44:57 Numerics, verification, and hardware that keeps changing
46:07 Are LLMs creative, or just good at search?
49:50 Getting models to extrapolate instead of interpolate
52:03 Why did we ever call it pre-training?
55:02 What RL is really learning
56:27 Competing with the big labs with fewer people
58:31 Will kernel generation keep old GPUs alive? Amdahl's law
1:01:58 Open source plans
1:02:55 Audience question: agentic optimizers
1:04:53 Audience question: Muon, Shampoo, and the future of second-order methods
1:08:58 Hiring at Core Automation


key topics

  • Journey from Google Brain to startup

  • Evolution of AI research and optimization

  • Pre-training and reinforcement learning

  • Kernel optimization and system efficiency

  • Open source AI and collaborative research

  • Challenges in AI creativity and exploration

  • Future directions in continual learning and model scaling


  • Music

  • "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0

Discussion about this episode

User's avatar

Ready for more?