用SAM2让机器人从不同角度都能准确抓取物体
Object-Centric Mobile Manipulation through SAM2-Guided Perception and Imitation Learning
- 基于SAM2的物体中心感知,融合抓取方向信息
- 在多角度演示下仍保持良好泛化能力,优于基线模型
- 适合需要灵活抓取的通用机器人系统开发
移动操作中的模仿学习是机器人操作领域的重要挑战。现有框架通常将导航与操作分离,仅在到达指定位置后执行操作,当导航不精确时易因接近角度错位导致性能下降。为使移动操作器能从不同朝向完成相同任务,构建通用机器人模型的关键能力,我们提出一种基于SAM2(支持可提示图像分割的基础模型)的物体中心方法,将操作朝向信息融入模型中,实现对同一任务在不同视角下的稳定理解。我们在自研移动操作机器人上部署该模型,并在多种朝向角度下评估其在抓取-放置任务上的表现。相较于动作分块变压器(Action Chunking Transformer),本模型在使用多角度示范训练时展现出更优的泛化能力。该工作显著提升了基于模仿学习的移动操作系统的泛化性与鲁棒性。
原文摘要 · Abstract (English)
Imitation learning for mobile manipulation is a key challenge in the field of robotic manipulation. However, current mobile manipulation frameworks typically decouple navigation and manipulation, executing manipulation only after reaching a certain location. This can lead to performance degradation when navigation is imprecise, especially due to misalignment in approach angles. To enable a mobile manipulator to perform the same task from diverse orientations, an essential capability for building general-purpose robotic models, we propose an object-centric method based on SAM2, a foundation model towards solving promptable visual segmentation in images, which incorporates manipulation orientation information into our model. Our approach enables consistent understanding of the same task from different orientations. We deploy the model on a custom-built mobile manipulator and evaluate it on a pick-and-place task under varied orientation angles. Compared to Action Chunking Transformer, our model maintains superior generalization when trained with demonstrations from varied approach angles. This work significantly enhances the generalization and robustness of imitation learning-based mobile manipulation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。