Tabular data is still where most of machine learning actually happens in industry, and the field has changed a lot in the last few years. In this episode we talk with David Holzmüller, a researcher at INRIA and one of the people behind TabArena, TabICL and RealMLP, about what the state of the art looks like right now and how to pick a model for your own data.
We cover the shift to TabPFN-style foundation models that learn to learn from whole tables, why TabArena was built and what earlier benchmarks got wrong, what Google's new TabFM means for the leaderboard, and when gradient boosted trees are still the right tool. David explains why LLMs struggle with tables, shares an early result comparing Claude Opus against TabICL on tiny datasets, and walks through how to embed text columns for tabular models. We also get into time series vs tabular data, the open research problems he thinks matter most, and why classical ML libraries are so bad out of the box.
Links:
TabArena: https://tabarena.ai
Topics
Tabular foundation models and in-context learning on tables
TabArena and Beyond Arena: building a benchmark that stays honest
TabFM, TabPFN, TabICL and the tradeoffs between them
When boosted trees and MLPs still win (large data, CPU, fast inference)
Why LLMs are inefficient on tabular data and where they might help
Embedding text columns with language models
Explainability, calibration and class imbalance
Time series vs tabular data
Open problems: invariances, synthetic data, uncertainty, scaling down
Where the field is heading in the next five years
Chapters
0:00 Intro
0:31 What changed in tabular ML: TabPFN-style foundation models
2:22 Which model to try first? TabArena and how it was built
5:14 What older benchmarks got wrong, and Beyond Arena
8:45 GPU AutoML vs foundation models
10:40 Reading the leaderboard: TabFM, TabPFN, TabICL and the tradeoffs
12:47 Calibration, class imbalance and small vs large data
19:45 Explainability for black-box tabular models
21:34 Why LLMs are bad at tabular data
25:39 Claude Opus 4.6 vs TabICL on tiny datasets
27:51 New classifiers, five-year outlook, real vs synthetic pretraining
33:13 Embedding text columns for tabular foundation models
36:13 Time series vs tabular data
39:59 When gradient boosted trees still win, and feature engineering
45:31 Open research problems and where the field is heading
52:54 Better MLPs and why classical defaults are bad out of the box
Music
"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0












