用可控制的世界模型模拟机器人操作,加速策略评估与优化。
Ctrl-World: A Controllable Generative World Model for Robot Manipulation

- 基于多视角和动作条件的可控世界模型,支持长时序交互。
- 在新场景下生成超20秒一致轨迹,性能排名准确率达95%以上。
- 适合机器人策略训练、仿真评估与自动化改进的研究者。
通用机器人策略虽能执行多种操作技能,但对陌生物体和指令的评估与改进仍具挑战。真实世界试错成本高,需大量数据与专家标注。世界模型提供可扩展的替代方案,在想象空间中进行策略回放。然而,现有模型难以支持与通用策略兼容的多步交互,缺乏多视角预测、细粒度动作控制和长期一致性。本文提出一个可控的多视角世界模型,通过姿态条件记忆检索机制保持长时序一致性,并以帧级动作条件实现精准控制。在包含95,000条轨迹、564个场景的DROID数据集上训练后,模型可在新场景和新相机位姿下生成超过20秒的空间-时间一致轨迹。实验表明,该方法无需真实机器人试运行即可准确评估策略性能;通过在想象中合成成功轨迹并用于监督微调,使策略成功率提升44.7%。
原文摘要 · Abstract (English)
Generalist robot policies can now perform a wide range of manipulation skills, but evaluating and improving their ability with unfamiliar objects and instructions remains a significant challenge. Rigorous evaluation requires a large number of real-world rollouts, while systematic improvement demands additional corrective data with expert labels. Both of these processes are slow, costly, and difficult to scale. World models offer a promising, scalable alternative by enabling policies to rollout within imagination space. However, a key challenge is building a controllable world model that can handle multi-step interactions with generalist robot policies. This requires a world model compatible with modern generalist policies by supporting multi-view prediction, fine-grained action control, and consistent long-horizon interactions, which is not achieved by previous works. In this paper, we make a step forward by introducing a controllable multi-view world model that can be used to evaluate and improve the instruction-following ability of generalist robot policies. Our model maintains long-horizon consistency with a pose-conditioned memory retrieval mechanism and achieves precise action control through frame-level action conditioning. Trained on the DROID dataset (95k trajectories, 564 scenes), our model generates spatially and temporally consistent trajectories under novel scenarios and new camera placements for over 20 seconds. We show that our method can accurately rank policy performance without real-world robot rollouts. Moreover, by synthesizing successful trajectories in imagination and using them for supervised fine-tuning, our approach can improve policy success by 44.7\%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。