Dhruv Batra spent years leading Embodied AI at Meta, training virtual robots to navigate photorealistic 3D scans of real buildings with pure reinforcement learning. Then he left to co-found Yutori and build agents for a very different environment: the web browser.
In this episode, Dhruv explains why he sees these as the same problem. Web agents, in his framing, are robots that act in a browser (pixels in, actions out), and the web turns out to be just as messy an environment as the physical world.
Along the way, we cover his definition of intelligence as “navigation in idea space,” why robotics is lagging LLMs, the sim-to-real gap and why you can’t fake friction coefficients, the teleoperation counterexample to the “it’s a sensor problem” argument, and his provocative claim that under the current paradigm, we solved machine learning and didn’t even realize it. He also makes the case for why the scaling hypothesis isn’t falsifiable, why JEPA-style arguments deserve to be grappled with, how Yutori trains its Navigator models with RL on live websites, and what happens to the ad-supported web when agents, not eyeballs, do the browsing.
Timeline
00:01 — Intro
00:54 — What embodied AI actually means
06:47 — Intelligence as navigation in idea space
13:26 — Habitat: training robots with pure RL, no maps
20:04 — Why robotics is behind LLMs
28:24 — Sim-to-real: what you can and can’t fake
33:34 — “We solved ML and nobody noticed”
37:12 — Leaving Meta, founding Yutori
43:21 — Web agents: screenshots in, actions out
48:15 — Why the web won’t rebuild itself for agents
53:32 — Training Navigator: RL on live websites
1:01:04 — Who pays for the web when agents browse?
1:09:17 — What Yutori means, closing thoughts
Music
“Kid Kodi” - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.
About
The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.












