The Information Bottleneck
The Information Bottleneck
Why Deep Learning Finally Works on Tables | Frank Hutter (Prior Labs)
0:00
-1:18:06

Why Deep Learning Finally Works on Tables | Frank Hutter (Prior Labs)

In this episode, Frank Hutter joins us to talk about TabPFN and why tabular data is suddenly the hottest problem in deep learning. Frank is a professor at the University of Freiburg and spent 15 years building the AutoML field before founding Prior Labs, which SAP just acquired for over a billion dollars.

We get into why deep learning failed on tables for a decade and what in-context learning changed, how TabPFN is trained entirely on synthetic data, and why a model that never saw a real time series ended up beating specialized forecasting models. Frank also explains the architecture tricks behind scaling from 10,000 to a million rows, where LLMs fit into data science (and where they embarrassingly don't), and what happens to XGBoost from here.

Beyond the research, Frank talks about the jump from professor to co-CEO, why he refused to merge his 45-person team into SAP's 110,000 employees, the open-weights licensing debate, and the case for building a frontier lab in Freiburg rather than San Francisco.


key topics

  • The role of foundation models in tabular data

  • Impact of SAP acquisition on Pro Labs

  • The evolution of AutoML and hyperparameter optimization

  • Challenges and solutions for large context in models

  • Open source models and licensing strategies

  • The importance of independence for startup agility

  • Future directions in AI for science and medicine


  • 00:00 Intro
    00:34 The SAP acquisition and staying independent
    07:39 Why tabular data is the next big thing in deep learning
    14:19 What makes tabular data hard
    19:14 AutoML, AutoGluon, and fifteen years of hyperparameter tuning
    28:27 Scaling TabPFN: context limits and architectures
    34:35 Agentic data science and LLMs
    39:30 Online learning, time series, and Bayesian inference in a forward pass
    47:05 Open weights and the license debate
    54:51 Will LLMs and tabular models merge?
    1:00:01 From academia to startup
    1:09:42 Why build in Europe
    1:12:53 Audience questions and hiring


  • Music

    • "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.

Discussion about this episode

User's avatar

Ready for more?