梳理机器人操作中的世界模型,理清预测什么、如何与动作结合、何时使用。
World Models for Robotic Manipulation: A Survey

- 按预测内容将世界模型分为五类,明确其与感知、策略的区别。
- 提出预测-动作连接的分类框架,区分集成模型与显式规划器。
- 适合研究机器人学习、具身智能和模拟器构建的开发者参考。
机器人操作依赖于在执行前预判动作如何改变物体、接触关系和场景几何。学习型世界模型通过预测任务相关的未来演化来实现这一能力,但该术语如今涵盖隐空间动力学模型、动作条件视频生成器、三维/四维场景预测器、物理启发模拟器以及视觉-语言-动作系统中的预测模块。这种广泛性导致文献碎片化,掩盖了对操作关键的设计选择。本文通过三个问题展开综述:预测的未来表征是什么?预测如何关联动作?预测在机器人学习流程中何时使用?我们操作性定义世界模型为动作条件预测系统,将其与感知模块、逆模型、策略、奖励函数和价值函数区分开。随后将现有工作归入五类表征家族,构建功能分类体系,分离集成预测-动作模型与显式预测规划器,并阐明合成经验生成、候选筛选、基于搜索的评估、学习环境和结果验证等基础设施角色。进一步将这些角色映射至预训练、后训练和推理适应阶段,回顾34个操作数据集,综合预测保真度、任务性能和模拟器可靠性评估协议。结果显示,世界模型正从任务特定的动力学预测器演变为机器人学习的预测基础设施,同时暴露接触建模、幻觉控制、动作对齐和闭环使用下的基准测试等开放挑战。
原文摘要 · Abstract (English)
Robotic manipulation depends on the ability to anticipate how actions reshape objects, contacts, and scene geometry before execution. Learned world models provide this capability by predicting task-relevant future evolution under robot intervention, yet the term now spans latent dynamics models, action-conditioned video generators, three- and four-dimensional scene predictors, physics-informed simulators, and predictive modules inside vision-language-action systems. This breadth has fragmented the literature and obscured the design choices that matter for manipulation. We survey world models for robotic manipulation through three questions: what future representation is predicted, how prediction is connected to action, and when prediction is used in the robot-learning pipeline. We operationally define a world model as an action-conditioned predictive system and distinguish it from perception modules, inverse models, policies, rewards, and value functions. We then organize existing work into five representation families, develop a functional taxonomy that separates integrated prediction-action models from explicit predictive planners, and characterize infrastructure roles including synthetic experience generation, candidate filtering, search-based evaluation, learned environments, and outcome verification. We further map these roles across pretraining, post-training, and inference adaptation, review 34 manipulation datasets, and synthesize evaluation protocols for predictive fidelity, task performance, and simulator reliability. This survey shows that world models are evolving from task-specific dynamics predictors into predictive infrastructure for robot learning, while exposing open challenges in contact modeling, hallucination control, action alignment, and benchmarking under closed-loop use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。