让机器人从具体物体学会通用操作,跨实例也能用
S$^2$-Diffusion: Generalizing from Instance-level to Category-level Skills in Robot Manipulation
- 用语义+空间联合建模捕捉动作本质
- 单张RGB图+深度网络实现类别泛化
- 真实世界与仿真测试均表现稳定
近期技能学习进展使机器人能从少量示范中掌握复杂操作,但这些技能通常仅限于训练中出现的具体物体和环境实例,难以迁移到同类其他实例。本文提出一种开放词汇的时空扩散策略(S²-Diffusion),实现从实例级到类别级的泛化,使技能可在同一类别的不同实例间迁移。通过可提示的语义模块结合空间表征,捕获技能的功能性特征;进一步利用深度估计网络,仅需单个RGB摄像头即可实现。在多种模拟与真实世界的机器人操作任务中进行评估,结果表明,S²-Diffusion对类别无关因素变化具有不变性,并能在未训练过的同类实例上保持良好性能。
原文摘要 · Abstract (English)
Recent advances in skill learning has propelled robot manipulation to new heights by enabling it to learn complex manipulation tasks from a practical number of demonstrations. However, these skills are often limited to the particular action, object, and environment \textit{instances} that are shown in the training data, and have trouble transferring to other instances of the same category. In this work we present an open-vocabulary Spatial-Semantic Diffusion policy (S$^2$-Diffusion) which enables generalization from instance-level training data to category-level, enabling skills to be transferable between instances of the same category. We show that functional aspects of skills can be captured via a promptable semantic module combined with a spatial representation. We further propose leveraging depth estimation networks to allow the use of only a single RGB camera. Our approach is evaluated and compared on a diverse number of robot manipulation tasks, both in simulation and in the real world. Our results show that S$^2$-Diffusion is invariant to changes in category-irrelevant factors as well as enables satisfying performance on other instances within the same category, even if it was not trained on that specific instance. Project website: https://s2-diffusion.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。