系统梳理世界模型的架构、方法与应用,助力智能体自主推理与规划。
World Models: A Comprehensive Survey of Architectures, Methodologies, Reasoning Paradigms, and Applications

- 构建四维分类框架:架构、方法族、推理策略、应用领域
- 整合从PlaNet到Sora等关键模型,揭示想象与思维链的融合趋势
- 适合研究通用人工智能、机器人与自动驾驶的开发者参考
世界模型作为内部模拟器,通过学习环境结构与动态规律,在追求通用人工智能中扮演核心角色,使智能体能在已学表征中进行预测、规划与推理。尽管在强化学习、机器人、自动驾驶和视频生成等领域进展迅速,但该领域仍缺乏统一框架来整合多样化的架构选择、训练方法、推理机制与应用场景。本文提出多维度分类体系,涵盖四方面:(i) 架构,包括表示形式、动态建模、输入模态、学习范式及下游应用;(ii) 方法族,如状态空间与递归模型、Transformer 模型、扩散生成器、物理信息网络、语言增强多模态系统;(iii) 推理策略,包括基于想象的规划、隐式策略学习、反事实推理与不确定性下的规划;(iv) 应用领域,覆盖机器人、自动驾驶、视频预测、多模态智能体、强化学习、科学建模、医学影像、教育评估及金融业务。从认知科学基础追溯至PlaNet、Dreamer系列、MuZero、Sora、Cosmos、Genie等里程碑系统,分析各维度交互关系,强调思维链推理与世界模型想象的近期融合。综述评估协议与基准,指出持续挑战如预测误差累积、仿真到现实迁移、评价碎片化,并展望统一多模态世界模型、规模化交互模拟器及高安全场景部署的未来方向。
原文摘要 · Abstract (English)
World models, internal simulators that learn the structure and dynamics of an environment, have emerged as a central paradigm in the pursuit of artificial general intelligence, enabling agents to predict, plan, and reason within learned representations. Despite rapid progress across reinforcement learning, robotics, autonomous driving, and video generation, the field lacks a unified framework integrating its diverse architectural choices, training methods, reasoning mechanisms, and application settings. This survey addresses that gap with a multi-axis taxonomy organized along four dimensions: (i) architecture, encompassing representation format, dynamics formulation, input modality, learning paradigm, and downstream application; (ii) methodological family, including state-space and recurrent approaches, transformer-based models, diffusion-based generators, physics-informed networks, and language-augmented multimodal systems; (iii) reasoning strategy, covering imagination-based planning, latent policy learning, counterfactual reasoning, and planning under uncertainty; and (iv) application domain, spanning robotics, autonomous driving, video prediction, multimodal agents, reinforcement learning, scientific modeling, medical imaging, educational measurement, and business and finance. Tracing the field from early cognitive-science foundations to milestone systems such as PlaNet, the Dreamer family, MuZero, Sora, Cosmos, and Genie, we examine how these dimensions interact and highlight the recent convergence of chain-of-thought reasoning with world-model imagination. We review evaluation protocols and benchmarks, identify persistent challenges such as compounding prediction errors, sim-to-real transfer, and fragmented evaluation, and outline future directions toward unified multimodal world models, foundation-scale interactive simulators, and safe deployment in safety-critical domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。