Should AI Learn to Dream to Achieve AGI?
From LLMs to World Models, the Race to go Beyond Language.
Image credits: Runway
This autumn, Yann LeCun, Meta’s chief AI scientist, left the company to launch a new startup, and the not-so-new subject of virtual worlds seemed to re-surface. Indeed, before this announcement, some major players were already making moves in the world-model direction. Just in August 2025, Google launched Genie 3, described as a “general purpose world model”, and about a year before, AI pioneer Fei-Fei Li raised $230 million for her new venture, World Labs, with the objective to build “frontier world models that can perceive, generate, reason, and interact with the 3D world” and launched its first commercial product Marble this November. Why this resurgence?
Artificial intelligence dominated tech news for the past few years, and our Digital Disruption Matrix ranked it 2025’s most disruptive technologies (descriptive & generative AI). The democratization of Large Language Models (LLMs) with the ChatGPT phenomenon did influence a growing number of objects, services and aspects of our lives. Yet, several of the architects of modern AI, including LeCun and Li, believe that mastering language is only one piece of the puzzle. They argue that world models represent the next crucial step towards truly intelligent systems.
What exactly is a World Model?
When World Labs AI engineer Christoph Lassner said “what if we could create an entire world where you can explore every angle and you can look around every corner” during this TedTalk in November 2025, it certainly brought back memories from the metaverse days but world models do differ from user facing virtual reality worlds and are certainly a critical evolution of LLMs.
Back in the 1940s, Kenneth Craik, a pioneer of modern cognitive science, theorized that an intelligent organism must carry a “small-scale model of external reality“ in its head and that the mind forms models of reality to predict future events. While LLMs are”book smart” by design, they lack the “street smarts” that come from a deep, intuitive understanding of physical reality. This gap between LLMs’ abstract knowledge and physical comprehension is the central problem that world models aim to solve. This also explains how the concept of a world model differs from the metaverse primarily in its function as an internal cognitive engine, rather than a platform for human interaction.
But as Yann LeCun puts it: “understanding the physical world is much more difficult than understanding the language.” He reinforces this by pointing out that despite our advances, we still “don’t have a robot capable of doing the same thing as a 5 or 6 year old child.” An LLM can write an essay about how to clear a dinner table, but it cannot yet power a robot to perform the task reliably. Their knowledge, derived entirely from text, is “ungrounded”. For Dr. Fei-Fei Li, “spatial intelligence” is key. She calls it the “scaffolding upon which our cognition is built.” And argues that this intuitive knowledge is used in everyday tasks, like parking a car but also for scientific breakthroughs. The goal with world models is for AI to learn about the world in the same way a human child does, not by being programmed explicitly with every rule for every situation but through observation and interaction.
For example, presented in Mastering Diverse Control Tasks through World Models, the algorithm DreamerV3 learns by building an internal world model and using it to “dream” future outcomes, enabling learning and planning without extensive real-world trial and error. This allowed it to become the first AI to collect diamonds in Minecraft entirely from scratch, without human demonstrations. With the same unchanged system, DreamerV3 also mastered over 150 diverse tasks, from Atari games to robotic control, highlighting world models as a way towards general intelligence.
Different approaches
The core mechanism of a world model is simulation. It allows an AI to ask “what if?” without acting in the real world. As Kenneth Craik envisioned, an internal model allows an organism to “try out various alternatives, conclude which is the best of them...and in every way to react in a much fuller, safer, and more competent manner.” However, the research community is currently split on how to achieve this, especially regarding whether intelligence requires simulating a super-photorealistic surface of reality or its abstract underlying rules.
On the photorealistic side, we have models like OpenAI’s Sora, Runway’s Gen-4.5, and Google DeepMind’s Genie posits that scaling video generation allows systems to implicitly learn physical laws by predicting future frames. These models aim to function as general-purpose simulators where agents can be trained in diverse virtual environments.
Image credits: Sima 2, Google DeepMind
World Labs focuses on “Spatial Intelligence” by generating persistent and editable 3D worlds rather than ephemeral photorealistic video clips, ensuring the consistent geometry and physics crucial for robotics training.
In contrast, Yann Le Cun and Meta advocate for the Joint Embedding Predictive Architecture (JEPA). Instead of attempting to reconstruct every detail of sensory input, this approach learns by predicting future states within an abstract representation space. JEPA relies on energy-based models to compare and predict high-level representations rather than raw pixels or signals. By doing so, it prioritizes learning stable and meaningful structures of the world, such as causal relationships, that are essential for planning and reasoning. This design allows computational resources to be focused on understanding how the world works, rather than on visually reproducing it in detail.
Finally, in a paper Critiques of World Models from July 2025, the authors argue that the primary purpose of a world model is not merely video generation, but to function as a “sandbox for reasoning” by simulating all actionable possibilities for purposeful decision-making, an ability they link to the psychological concept of “hypothetical thinking.” The text systematically critiques several prevailing approaches and proposes a new framework, the Physical, Agentic, and Nested (PAN), that aims to bridge high-level reasoning with real-world physical sensations, allowing an agent to mentally rehearse complex actions before performing them.
Beyond these debates, the development of world models faces some important challenges, notably in terms of the demand for data and computational power, that is, just like for the AI sector in general, forming a bottleneck for innovation.
From empowering creatives to embodied AI
Image credits: What Is Embodied AI?, NVIDIA
Right now, commercial world models like Marble are mostly directed towards filmmakers, architects, and game designers, allowing them to easily generate and interactively edit complex 3D worlds with simple text or image prompts. SIMA 2 from Google DeepMind is also focused on video games, creating an interactive gaming companion. But for both World Labs and Google DeepMind, those are just steps towards the development of Artificial General Intelligence (AGI) and embodied AI. As stated on SIMA 2 blogpage: “This is a significant step in the direction of AGI, with important implications for the future of robotics and AI-embodiment in general”. The next step is with embodied AI also because real-world robot training data is scarce and dangerous to collect, they could use their world models as simulators to generate massive amounts of synthetic training data. Embodied AI is also the core focus of NVIDIA, which developed NVIDIA Cosmos and aims to develop world models for industrial and robotics applications, such as factory robots, warehouse automation, and autonomous vehicles.
As the subject of world models gains importance and conveys disruptive power for the creative industry, robotics, AI agents and more, we will discuss and analyse it in our 2026 Digital Disruption Matrix.
This week’s curated news:
Meta’s Yann LeCun targets €3bn valuation for new AI start-up
Yann LeCun is in early talks to raise €500m for Advanced Machine Intelligence Labs, a new AI company focused on advanced “world model” systems, marking his next move as he prepares to leave Meta after more than a decade.
Read more here
YouTube wins exclusive rights to stream the Oscars from 2029
YouTube will become the exclusive home of the Oscars starting in 2029, ending ABC’s decades-long run and marking a major shift as one of Hollywood’s biggest live events moves fully from broadcast TV to streaming.
Read more here
Amazon in talks to invest $10bn in OpenAI at $500bn-plus valuation
Amazon is reportedly in early discussions to invest up to $10bn in OpenAI, a deal that could value the AI company at more than $500bn and deepen ties around cloud infrastructure and AI chips as “circular deals” reshape the AI ecosystem.
Read more here
Bezos and Musk race to take data centers into space
Jeff Bezos’ Blue Origin and Elon Musk’s SpaceX are exploring orbital data centers to support AI computing, extending their long-running rivalry as both bet on space-based infrastructure to tap into the booming demand for AI capacity.
Read more here
🔍 Dive Deeper into Digital Disruption
Explore the 2025 Digital Disruption Matrix → Your guide to navigating digital transformation with confidence. This annual barometer combines rigorous data analysis with human perspective to rank this year’s most disruptive technologies. Discover how blockchain, AI, Web3, and emerging innovations are reshaping industries.
💡 Never Miss an Insight
To get weekly analysis on emerging technologies and digital transformation delivered to your inbox, follow us on Substack.
Follow the Digital Disruption Chair on LinkedIn for regular insights, interviews, comments and articles on frontier technologies.
Learn more about our programs | Contact us
Bibliography:
a16z (Director). (2024, September 20). “The Future of AI is Here”—Fei-Fei Li Unveils the Next Frontier of AI [Video recording], Youtube.
About Runway AI Video Platform | Runway AI. (n.d.). Retrieved December 18, 2025, from https://runwayml.com/about
Bellan, I. M., Rebecca. (2025, December 11). Runway releases its first world model, adds native audio to latest video model. TechCrunch. https://techcrunch.com/2025/12/11/runway-releases-its-first-world-model-adds-native-audio-to-latest-video-model/
Bellan, R. (2025, November 12). Fei-Fei Li’s World Labs speeds up the world model race with Marble, its first commercial product. TechCrunch. https://techcrunch.com/2025/11/12/fei-fei-lis-world-labs-speeds-up-the-world-model-race-with-marble-its-first-commercial-product/
Bilawal Sidhu (Director). (2025, November 15). This Changes Games Forever: AI That Builds and Plays in 3D Worlds[Video recording], Youtube.
Boone, J. (2025, December 5). « Nous n’avons pas de robot capable de faire la même chose qu’un enfant de 5 ou 6 ans »: Les « world models », nouvelle frontière de l’IA. Les Echos. https://www.lesechos.fr/tech-medias/intelligence-artificielle/nous-navons-pas-de-robot-capable-de-faire-la-meme-chose-quun-enfant-de-5-ou-6-ans-les-world-models-nouvelle-frontiere-de-lia-2202886
Del Ser, J., Lobo, J. L., Müller, H., & Holzinger, A. (2025). World Models in Artificial Intelligence: Sensing, Learning, and Reasoning Like a Child (Version 1). arXiv. https://doi.org/10.48550/ARXIV.2503.15168
Ding, J., Zhang, Y., Shang, Y., Feng, J., Zhang, Y., Zong, Z., Yuan, Y., Su, H., Li, N., Piao, J., Deng, Y., Sukiennik, N., Gao, C., Xu, F., & Li, Y. (2025). Understanding World or Predicting Future? A Comprehensive Survey of World Models (No. arXiv:2411.14499). arXiv. https://doi.org/10.48550/arXiv.2411.14499
Genie 3: A new frontier for world models. (n.d.). Google DeepMind. Retrieved December 18, 2025, from https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models/
Germanidis. (2023, December 11). Runway Research | Introducing General World Models. https://runwayml.com/research/introducing-general-world-models
Gupta, T., & Pruthi, D. (2025). Beyond World Models: Rethinking Understanding in AI Models (No. arXiv:2511.12239; Version 1). arXiv. https://doi.org/10.48550/arXiv.2511.12239
Hafner, D., Pasukonis, J., Ba, J., & Lillicrap, T. (2025). Mastering diverse control tasks through world models. Nature, 640(8059), 647–653. https://doi.org/10.1038/s41586-025-08744-2
Harvard CMSA (Director). (2025, September 29). Yann LeCun | Self-Supervised Learning, JEPA, World Models, and the future of AI [Video recording], Youtube.
Hendrycks, D., Song, D., Szegedy, C., Lee, H., Gal, Y., Brynjolfsson, E., Li, S., Zou, A., Levine, L., Han, B., Fu, J., Liu, Z., Shin, J., Lee, K., Mazeika, M., Phan, L., Ingebretsen, G., Khoja, A., Xie, C., … Bengio, Y. (2025). A Definition of AGI (Version 3). arXiv. https://doi.org/10.48550/ARXIV.2510.18212
Latent Space (Director). (2025, November 25). After LLMs: Spatial Intelligence and World Models — Fei-Fei Li & Justin Johnson, World Labs [Video recording], Youtube.
LeCun, Y. (2022). A Path Towards Autonomous Machine Intelligence Version 0.9.2, 2022-06-27. Courant Institute of Mathematical Sciences, New York University, Meta - Fundamental AI Research.
Lenny’s Podcast (Director). (2025, November 16). The Godmother of AI on jobs, robots & why world models are next | Dr. Fei-Fei Li [Video recording], Youtube.
Li, F.-F. (2025, November 10). From Words to Worlds: Spatial Intelligence is AI’s Next Frontier [Substack newsletter]. Dr. Fei-Fei Li.
Marble: A Multimodal World Model. (n.d.). Retrieved December 18, 2025, from https://www.worldlabs.ai/blog/marble-world-model
Mims, C. (2025, September 26). What Are ‘World Models’? The Key to the Next Big AI Leap. Wall Street Journal. https://www.wsj.com/tech/ai/world-models-ai-evolution-11275913
Pavlus, J. (2025, September 2). ‘World Models,’ an Old Idea in AI, Mount a Comeback. Quanta Magazine. https://www.quantamagazine.org/world-models-an-old-idea-in-ai-mount-a-comeback-20250902/
Recherches—Des recherches de pointe avec l’AGI en ligne de mire. (2025, November 12). https://openai.com/fr-FR/research/
Research—We work on some of the most complex and interesting challenges in AI. (2025, November). Google DeepMind. https://deepmind.google/research/
Scaling Robotic Simulation with Marble. (2025, November 12). https://www.worldlabs.ai/case-studies/1-robotics
SIMA 2: A Gemini-Powered AI Agent for 3D Virtual Worlds. (2025, August). Google DeepMind. https://deepmind.google/blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worlds/
What Are World Models and How Are They Built? (n.d.). NVIDIA. Retrieved December 18, 2025, from https://www.nvidia.com/en-us/glossary/world-models/
Wiggers, K. (2024, December 14). What are AI “world models,” and why do they matter? TechCrunch. https://techcrunch.com/2024/12/14/what-are-ai-world-models-and-why-do-they-matter/
World Labs (Director). (2025, November 13). World Labs: Just Imagine [Video recording], Youtube.
Xing, E., Deng, M., Hou, J., & Hu, Z. (2025). Critiques of World Models (No. arXiv:2507.05169). arXiv. https://doi.org/10.48550/arXiv.2507.05169



