Andrew Dai spent over a decade at Google Brain and DeepMind, where he co-wrote the 2015 paper that introduced language model pre-training followed by fine-tuning, and later co-led pre-training data for Gemini. He's now co-founder and CEO of Elorian AI, which is building models for visual reasoning.
We talk about how his pre-training result started as a bug, why next-token prediction scales better than other objectives, and what went wrong for Google in the early LLM race. The second half is about vision: why today's frontier models still can't count objects in a photo, why he thinks reasoning is fundamentally visual, and how his view of world models differs from JEPA.
Chapters
00:00 Intro
00:48 Andrew's background and the accidental discovery of pre-training
07:10 Why next-token prediction scales
10:09 How Google fell behind and the early days of Gemini
18:28 What makes training data good
28:16 Where visual understanding breaks down
46:01 World models, JEPA and robotics
55:27 Generation vs understanding, and what's next for Elorian
Topics
Pre-training and fine-tuning
Scaling and next-token prediction
Gemini and Google's AI history
Data quality and synthetic data
Visual reasoning and counting
World models and JEPA
Music
"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0









