Most labs build language models by scraping the web and filtering afterward. Pierre-Carl Langlais runs it the other way around. At Pleias, the French-German lab he co-founded, the models are built from data he can actually account for, which in practice means open and public-domain sources plus a lot of synthetic data the lab generates itself. It sounds like a self-imposed handicap. It mostly isn’t. One of their models is a 600 million parameter system that runs live inside the Paris subway’s monitoring pipeline.
We cover the SYNTH pretraining dataset and why he thinks “ethical data” has to mean more than copyright-free. He explains why barely 2% of their Common Corpus appears in typical web crawls, and why that gap is really a preservation problem. From there, he gets blunt about benchmark maxing and whether GLM really earns its Opus-class reputation. He also argues that the quiet move by closed labs to hide reasoning traces is mostly about claiming ownership of model outputs. He’s skeptical of sovereign AI, and not shy about how Mistral drifted from frontier research toward French corporate consulting. We finish on NVIDIA’s persona datasets and the odd idea of training on the conditions that produced a text rather than the text itself.
Timeline
(00:02) Welcome and introductions
(00:49) Why synthetic data matters, and the SYNTH set
(04:15) Three reasons to control your training data
(07:18) What “ethical data” actually means
(11:08) How Common Corpus got built, from Wikipedia to PDFs
(16:35) Agentic harnesses and synthetic data
(20:03) Evaluating data when you train on reasoning traces
(25:27) General versus specialized pretraining
(27:08) Benchmark maxing and the GLM question
(31:51) Getting diversity in, and the NVIDIA personas
(35:02) Hidden reasoning traces and the fight over model IP
(38:17) Mid-training and the “It’s All Training” thesis
(41:47) Can small models actually compete
(45:01) Cybersecurity and Europe’s strategic gap
(47:08) Do you need a big model to orchestrate the small ones
(52:08) Sovereign AI and the limits of national champions
(56:42) Scaling laws when you control the data
(01:00:41) The NVIDIA persona datasets
(01:04:52) What you actually do with synthetic personas
(01:08:22) Closing thoughts
Music
“Kid Kodi” - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.
About
The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.












