系统梳理具身智能世界模型,构建分类框架并指出关键挑战。
A Comprehensive Survey on World Models for Embodied AI
- 提出三轴分类法:功能、时序建模与空间表示
- 汇总机器人、自动驾驶等领域的数据与评估指标
- 适合研究具身智能与世界模型的学者参考
具身智能需要能感知、行动并预判行为如何改变未来世界状态的智能体。世界模型作为内部模拟器,捕捉环境动态,支持前向与反事实推演,以支撑感知、预测与决策。本文提出具身智能中世界模型的统一框架,形式化问题设定与学习目标,构建涵盖三方面的分类体系:(1)功能性,即决策耦合型与通用型;(2)时序建模,包括序列模拟与推理、全局差异预测;(3)空间表征,包括全局隐向量、标记特征序列、空间隐网格与分解渲染表示。系统整理机器人、自动驾驶及通用视频场景下的数据资源与评估指标,涵盖像素预测质量、状态理解水平与任务性能。进一步对前沿模型进行定量对比,并提炼核心开放挑战:统一数据集稀缺、评估需关注物理一致性而非仅像素保真度、模型性能与实时控制所需计算效率间的权衡,以及长时程时间一致性与误差累积抑制的建模难题。项目维护的精选文献库见 https://github.com/Li-Zn-H/AwesomeWorldModels。
原文摘要 · Abstract (English)
Embodied AI requires agents that perceive, act, and anticipate how actions reshape future world states. World models serve as internal simulators that capture environment dynamics, enabling forward and counterfactual rollouts to support perception, prediction, and decision making. This survey presents a unified framework for world models in embodied AI. Specifically, we formalize the problem setting and learning objectives, and propose a three-axis taxonomy encompassing: (1) Functionality, Decision-Coupled vs. General-Purpose; (2) Temporal Modeling, Sequential Simulation and Inference vs. Global Difference Prediction; (3) Spatial Representation, Global Latent Vector, Token Feature Sequence, Spatial Latent Grid, and Decomposed Rendering Representation. We systematize data resources and metrics across robotics, autonomous driving, and general video settings, covering pixel prediction quality, state-level understanding, and task performance. Furthermore, we offer a quantitative comparison of state-of-the-art models and distill key open challenges, including the scarcity of unified datasets and the need for evaluation metrics that assess physical consistency over pixel fidelity, the trade-off between model performance and the computational efficiency required for real-time control, and the core modeling difficulty of achieving long-horizon temporal consistency while mitigating error accumulation. Finally, we maintain a curated bibliography at https://github.com/Li-Zn-H/AwesomeWorldModels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。