让移动操作机器人学会从关键视角和区域观察,提升任务鲁棒性。
Robust Imitation Learning for Mobile Manipulator Focusing on Task-Related Viewpoints and Regions
- 通过多视角注意力机制学习任务相关视角与区域
- 在不同环境中成功率最高提升29.3点
- 适合需要抗遮挡和跨环境适应的机器人应用
本文研究移动操作机器人从视觉观测角度泛化视觉-运动策略的问题。当仅采用单一视角时,机器人因自身结构导致遮挡,部署于不同场景时又面临显著领域偏移。然而,据作者所知,尚无研究能同时解决遮挡与领域偏移问题并提出鲁棒策略。为此,本文提出一种面向任务相关视角及其空间区域的多视角模仿学习方法。该方法通过增强数据集训练的注意力机制,获得最优视角与对遮挡及领域偏移具有鲁棒性的视觉嵌入。在不同任务与环境下的对比实验表明,本方法成功率达29.3点提升。消融实验显示,从多视角数据中学习任务相关视角可增强抗遮挡能力;聚焦任务相关区域可使成功率在领域偏移下提升达33.3点。
原文摘要 · Abstract (English)
We study how to generalize the visuomotor policy of a mobile manipulator from the perspective of visual observations. The mobile manipulator is prone to occlusion owing to its own body when only a single viewpoint is employed and a significant domain shift when deployed in diverse situations. However, to the best of the authors' knowledge, no study has been able to solve occlusion and domain shift simultaneously and propose a robust policy. In this paper, we propose a robust imitation learning method for mobile manipulators that focuses on task-related viewpoints and their spatial regions when observing multiple viewpoints. The multiple viewpoint policy includes attention mechanism, which is learned with an augmented dataset, and brings optimal viewpoints and robust visual embedding against occlusion and domain shift. Comparison of our results for different tasks and environments with those of previous studies revealed that our proposed method improves the success rate by up to 29.3 points. We also conduct ablation studies using our proposed method. Learning task-related viewpoints from the multiple viewpoints dataset increases robustness to occlusion than using a uniquely defined viewpoint. Focusing on task-related regions contributes to up to a 33.3-point improvement in the success rate against domain shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。