用扩散模型生成可控的机器人操作视频,提升仿真到现实的迁移效果。
RoboTransfer: Controllable Geometry-Consistent Video Diffusion for Manipulation Policy Transfer
- 基于扩散模型生成多视角一致的机器人操作视频
- 合成数据训练的策略在未见场景中泛化性能更优
- 支持背景编辑和物体替换等细粒度控制,适合仿真实验
通用机器人目标是让智能体能无缝适应并操作于多样化的非结构化人类环境。模仿学习已成为机器人操作的关键范式,但大规模、多样化的示范数据收集成本过高。模拟器提供了一种低成本替代方案,但仿真到现实的差距仍是可扩展性的主要障碍。我们提出RoboTransfer,一种基于扩散模型的视频生成框架,用于合成机器人数据。通过利用跨视角特征交互和全局一致的3D几何信息,RoboTransfer在保证多视角几何一致性的同时,实现了对场景元素(如背景编辑、物体替换)的细粒度控制。大量实验表明,RoboTransfer生成的视频具有更优的几何一致性和视觉保真度。此外,在合成数据上训练的策略在面对新出现的未见场景时表现出更强的泛化能力。项目页面:https://horizonrobotics.github.io/robot_lab/robotransfer。
原文摘要 · Abstract (English)
The goal of general-purpose robotics is to create agents that can seamlessly adapt to and operate in diverse, unstructured human environments. Imitation learning has become a key paradigm for robotic manipulation, yet collecting large-scale and diverse demonstrations is prohibitively expensive. Simulators provide a cost-effective alternative, but the sim-to-real gap remains a major obstacle to scalability. We present RoboTransfer, a diffusion-based video generation framework for synthesizing robotic data. By leveraging cross-view feature interactions and globally consistent 3D geometry, RoboTransfer ensures multi-view geometric consistency while enabling fine-grained control over scene elements, such as background editing and object replacement. Extensive experiments demonstrate that RoboTransfer produces videos with superior geometric consistency and visual fidelity. Furthermore, policies trained on this synthetic data exhibit enhanced generalization to novel, unseen scenarios. Project page: https://horizonrobotics.github.io/robot_lab/robotransfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。