Florian Brand builds evals at Prime Intellect. The premise of the conversation is that writing a benchmark is the easy part now. Keeping the model from cheating it is the job, and it takes longer than the benchmark itself.
We get into why he thinks you can't evaluate a model apart from the CLI it runs in, what happens to statistics when a single run costs five figures, and whether the feeling that a model just works can ever become a number.
He also has a few stories about agents finding their way around the scoring that are worth hearing cold.
Timeline
00:13 Intro
01:00 What evals are for
04:05 Agentic benchmarks
07:10 Kimi K2 and model diversity
08:23 Long-horizon coding tasks
10:29 Building a benchmark
12:15 MirrorCode
14:27 Rubrics and LLM judges
16:30 The cost of expert labelers
17:49 Long runs and variance
19:44 Evaluating the harness
24:29 Chinese labs building CLIs
30:00 More reward hacking
37:45 Tau-bench and economic tasks
39:43 Benchmaxxing and GLM 5.2
45:15 Statistics and cost
47:56 Frontier convergence
52:04 Misuse in open and closed models
55:35 Self-improvement
Music
"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.
About
The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.









