In this episode, Frank Hutter joins us to talk about TabPFN and why tabular data is suddenly the hottest problem in deep learning. Frank is a professor at the University of Freiburg and spent 15 years building the AutoML field before founding Prior Labs, which SAP just acquired for over a billion dollars.
We get into why deep learning failed on tables for a decade and what in-context learning changed, how TabPFN is trained entirely on synthetic data, and why a model that never saw a real time series ended up beating specialized forecasting models. Frank also explains the architecture tricks behind scaling from 10,000 to a million rows, where LLMs fit into data science (and where they embarrassingly don't), and what happens to XGBoost from here.
Beyond the research, Frank talks about the jump from professor to co-CEO, why he refused to merge his 45-person team into SAP's 110,000 employees, the open-weights licensing debate, and the case for building a frontier lab in Freiburg rather than San Francisco.
key topics
The role of foundation models in tabular data
Impact of SAP acquisition on Pro Labs
The evolution of AutoML and hyperparameter optimization
Challenges and solutions for large context in models
Open source models and licensing strategies
The importance of independence for startup agility
Future directions in AI for science and medicine
00:00 Intro
00:34 The SAP acquisition and staying independent
07:39 Why tabular data is the next big thing in deep learning
14:19 What makes tabular data hard
19:14 AutoML, AutoGluon, and fifteen years of hyperparameter tuning
28:27 Scaling TabPFN: context limits and architectures
34:35 Agentic data science and LLMs
39:30 Online learning, time series, and Bayesian inference in a forward pass
47:05 Open weights and the license debate
54:51 Will LLMs and tabular models merge?
1:00:01 From academia to startup
1:09:42 Why build in Europe
1:12:53 Audience questions and hiring
Music
"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.











