Daphne Koller wrote the book that many of us learned probabilistic graphical models from, founded Coursera, and now runs insitro, which is trying to make drug discovery a machine-learning problem.
We start with the bitter lesson. She agrees with most of it and then says where it stops working: biology doesn’t have enough data, structure is how people understand anything, and making a drug is a question about an intervention that hasn’t happened yet, not a pattern in data you already have.
Most of the episode is about why drug discovery is hard. Ninety percent of drugs that reach the clinic fail, and mostly not because the molecule was bad. The molecule usually does what it was designed to do. It just turns out the thing it was designed to do had nothing to do with the disease. Only 22% of diseases have any approved drug at all, and she calls that an upper bound on what we understand, not a lower bound.
She also gets into what agents are and aren’t good for in a wet lab, why cells don’t grow faster no matter how many GPUs you point at them, what it would take to have real foundation models for biology, and why almost all of biology is still out of distribution.
Plus GLP-1s and what human data keeps teaching us, whether AI can make the kind of leap that turned a bacterial immune system into CRISPR, and what she’d build if she were starting Coursera today.
Key Topics
The impact of scaling and data in machine learning
The importance of structure and causality in AI
Challenges in drug discovery and biological understanding
The role of foundation models in biology
Ethical considerations in AI and biomedical research
Chapters
00:00 Introduction to Machine Learning and Drug Discovery
02:00 The Bitter Lesson and Its Implications
06:48 Challenges in Drug Design and Discovery
11:48 Ethical Considerations in Human Research
17:20 The Drug Discovery Pipeline Explained
29:30 Integrating AI in Experimental Design
35:38 The Role of Human Judgment in Drug Design
37:14 Future of Drug Design: Efficiency vs. Automation
39:37 Challenges in AI and Data Availability for Biology
41:08 Foundation Models: Potential and Limitations
43:39 Causality in Biological Data: Importance and Challenges
45:18 Creativity vs. Understanding in Drug Design
48:17 Balancing Investments in Data, Algorithms, and Experiments
50:07 The Value of Simulations in Drug Discovery
52:03 Mathematical Frameworks in Biology: Utility and Limitations
54:14 The Future of Drug Discovery: Optimism and Innovations
56:28 The Impact of Coursera on Education
01:00:33 The Role of Universities in Lifelong Learning
01:04:06 Connecting Dots: The Fun of Variety in Work
01:05:46 Optimism for the Future of Drug Discovery
Music
“Kid Kodi” - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.











