仅用一次真人示范,生成海量合成数据训练机器人在复杂环境中自主操作。
Worlds in One Demo: A Synthetic Data Engine for Learning Open-World Mobile Manipulation

- 基于一次演示重建环境与交互轨迹,通过运动规划重组生成新任务序列。
- 在模拟的多样化3D场景中合成数据,使策略具备长程鲁棒性和跨环境泛化能力。
- 支持不同机器人形态的零样本部署,适用于多场景、多设备的开放世界操作。
学习开放世界的移动操作策略需要大量数据以实现空间泛化、长时程鲁棒性及场景泛化。当前主流的数据采集方式(遥操作和UMI)在大规模下人力成本过高。为突破人工数据收集的瓶颈,我们提出WANDA:仅需一次人类示范,即可通过合成数据引擎学习开放世界移动操作。WANDA首先从源RGBD观测中重建背景高斯点云与机器人-物体交互轨迹,作为后续规划与渲染的世界基底;随后将富含接触的交互片段重新排列成多样化的空间配置,并利用全身体运动规划将其串联为新轨迹。为增强长时程鲁棒性,引入修正状态扩展以提升各阶段机器人与物体状态多样性;为实现跨环境泛化,所有轨迹在由日常照片生成的多样化3D世界中合成。此外,通过将渲染的机器人与物体网格与高斯点云背景拼接,生成逼真的观测数据。我们在多种仿真与真实场景的任务中评估该方法,实验表明,仅用一次真实示范训练出的策略即具备长时程鲁棒性、广泛空间泛化和跨环境泛化能力。同时,WANDA可自然支持跨机体数据生成,在另一具不同构型的移动操作器上实现零样本部署。
原文摘要 · Abstract (English)
Learning open-world mobile manipulation policies requires vast data to achieve spatial generalization, long-horizon robustness, and scene generalization. Current prevailing data collection paradigms, teleoperation and UMI, demand prohibitive human effort and cost at scale. To scale beyond the limits of manual data collection, we seek to maximize the value of each human demonstration by scalable data generation. To this end, we introduce WANDA: learning open-World mobile mANipulation from one demonstration via a synthetic DAta engine. WANDA first reconstructs background Gaussian splats and robot-object interaction trajectories from source RGBD observations, as a world substrate for later planning and rendering. It then rearranges contact-rich robot-object interaction segments into extensive spatial configurations, utilizing whole-body motion planning to chain them into new trajectories. To enhance long-horizon robustness, it applies Corrective State Expansion to increase the robot and object state diversity at different stages of mobile manipulation. To unlock cross-environment generalization, trajectories are synthesized on diverse generated 3D worlds from everyday photos. Furthermore, we synthesize photo-realistic observations by compositing rendered robot and object meshes with Gaussian splatting backgrounds. We evaluate our approach on extensive simulation and real-world tasks in various scenes. Experiments show that policies trained with WANDA achieve long-horizon robustness, broad spatial generalization and cross-environment generalization from one real demonstration. Moreover, WANDA naturally supports cross-embodiment data generation, validated by zero-shot deployment on another mobile manipulator with a distinct morphology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。