仅用10次示范就学会通用3D操作,还能适应不同位置和视角。
Learning Generalizable 3D Manipulation With 10 Demonstrations
- 通过语义引导感知与扩散决策模块,从10次演示中学习空间泛化技能。
- 在复杂变体下成功率达60%提升,优于现有最先进方法。
- 适合需要少量数据、高泛化能力的工业与服务机器人场景。
从示范中学习鲁棒且可泛化的操作技能仍是机器人领域的关键挑战,广泛应用于工业自动化与服务机器人。尽管近期模仿学习方法取得显著成果,但通常需要大量示范数据,且难以跨空间变化泛化。本文提出一种新框架,仅需10次示范即可学习操作技能,并泛化至不同初始物体位置与相机视角。框架包含两个核心模块:语义引导感知(SGP),从RGB-D输入构建任务导向、空间感知的3D点云表示;空间泛化决策(SGD),基于扩散模型生成动作。为在有限数据下学习泛化能力,引入关键的空间等变训练策略,捕捉专家示范中的空间知识。在仿真基准与真实机器人系统上进行大量实验验证,本方法在一系列挑战性任务中成功率较现有最优方法提升60%,即使面对物体姿态和相机视角的大幅变化仍表现优异。该工作展示了在真实应用中实现高效、可泛化操作技能学习的巨大潜力。
原文摘要 · Abstract (English)
Learning robust and generalizable manipulation skills from demonstrations remains a key challenge in robotics, with broad applications in industrial automation and service robotics. While recent imitation learning methods have achieved impressive results, they often require large amounts of demonstration data and struggle to generalize across different spatial variants. In this work, we present a novel framework that learns manipulation skills from as few as 10 demonstrations, yet still generalizes to spatial variants such as different initial object positions and camera viewpoints. Our framework consists of two key modules: Semantic Guided Perception (SGP), which constructs task-focused, spatially aware 3D point cloud representations from RGB-D inputs; and Spatial Generalized Decision (SGD), an efficient diffusion-based decision-making module that generates actions via denoising. To effectively learn generalization ability from limited data, we introduce a critical spatially equivariant training strategy that captures the spatial knowledge embedded in expert demonstrations. We validate our framework through extensive experiments on both simulation benchmarks and real-world robotic systems. Our method demonstrates a 60 percent improvement in success rates over state-of-the-art approaches on a series of challenging tasks, even with substantial variations in object poses and camera viewpoints. This work shows significant potential for advancing efficient, generalizable manipulation skill learning in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。