arXiv:2502.16932cs.RO2025-02被引 108

用合成演示提升机器人操作的少样本学习能力

DemoGen: Synthetic Demonstration Generation for Data-Efficient Visuomotor Policy Learning

  • 仅需1个真人示范,通过3D点云生成新场景下的动作演示
  • 在真实任务中显著提升策略性能,支持柔性物体等复杂场景
  • 可扩展实现抗干扰和避障等新能力,适合少数据场景

视觉-运动策略在机器人操作中展现出巨大潜力,但通常需要大量人类收集的数据才能有效运行。数据需求高的主要原因是其空间泛化能力有限,需在不同物体配置下广泛采样。本文提出DemoGen,一种低成本、完全合成的自动演示生成方法。每个任务仅需一个真人示范,便可通过将动作轨迹适配到新物体配置,生成空间增强的演示。视觉观察通过3D点云作为模态,并借助3D编辑重新排列场景中的物体来合成。实证表明,DemoGen在多种真实世界操作任务中显著提升策略性能,即使在涉及柔性物体、灵巧手末端执行器和双臂平台的挑战性场景中也适用。此外,DemoGen还可扩展以支持额外的分布外能力,包括抗扰动和障碍物避让。

原文摘要 · Abstract (English)

Visuomotor policies have shown great promise in robotic manipulation but often require substantial amounts of human-collected data for effective performance. A key reason underlying the data demands is their limited spatial generalization capability, which necessitates extensive data collection across different object configurations. In this work, we present DemoGen, a low-cost, fully synthetic approach for automatic demonstration generation. Using only one human-collected demonstration per task, DemoGen generates spatially augmented demonstrations by adapting the demonstrated action trajectory to novel object configurations. Visual observations are synthesized by leveraging 3D point clouds as the modality and rearranging the subjects in the scene via 3D editing. Empirically, DemoGen significantly enhances policy performance across a diverse range of real-world manipulation tasks, showing its applicability even in challenging scenarios involving deformable objects, dexterous hand end-effectors, and bimanual platforms. Furthermore, DemoGen can be extended to enable additional out-of-distribution capabilities, including disturbance resistance and obstacle avoidance.

机器人操作少样本学习合成数据3D生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。