world models
Notes and deep-dives on world models — learned simulators, predictive representations, and visual cognition for embodied agents.
World Models
Notes and deep-dives on world models — learned simulators, predictive representations, and visual cognition for embodied agents.
00 World Models — Part 0: From Language Models to World Models A ground-up introduction: why next-token prediction is not enough, what a world model actually is, and why learning to predict the future of an environment may be the next step toward grounded, agentic intelligence. 01 World Models — Part 1: Inside the Latent (VAEs, RSSM, and Learning in a Dream) The machinery, derived from scratch: why we compress pixels into a latent, how the VAE and the ELBO actually work, how a Recurrent State-Space Model turns a single-frame encoder into a simulator, and how Dreamer trains a policy entirely inside its own imagination — with the real systems that run on these ideas today.
Reading & resources
World Models — Reading & Resources: Deep Dives + Annotated Idea Map One place for the World Models literature: in-depth read-throughs of the key papers (gist, pipeline, results, stated future work, and idea openings) followed by a broad, thematically organized index. 🎯 marks the work closest to my active-perception research direction. Embodied IQA & Active View Selection — Deep Paper Review A full deep review of three papers (Embodied-IQA, EPD/MA-EIQA, Active View Selector): per-paper task/dataset/method/evaluation, cross-paper synthesis, a new-metric opportunity analysis, gaps and future directions, and a 2024–2026 related-work map. Embodied IQA — A Plain Guide: What the Papers Say, What Is Missing, and What I Want to Build A simple, jargon-light companion to the deep review. Each paper in four short sentences, a glossary of the abbreviations, the seven gaps that still block progress, and the research pipeline I plan to follow — from a noise-ceiling audit to a task-aware quality metric that tells a robot where to move.