让机器人视觉注意力可解释可修正,提升复杂场景下的操作鲁棒性
GuidedAttention: Interpretable and Correctable Visual Attention for OOD-Robust Robot Manipulation via Imitation Learning

- 通过预测关键点作为视觉注意力中间表示
- 在分布外条件下任务成功率提升显著
- 适合需要人机协作与安全调试的机器人应用
端到端视觉-运动策略难以让人类理解或修正其视觉关注点。我们提出GuidedAttention,一种引入可解释且可修正视觉注意力的视觉-运动模仿学习框架。从摄像头图像中预测与任务相关的注意力关键点,并以此条件化基于扩散模型的动作策略。用户可在推理初始化时检查并选择性修正关键点,随后由追踪模块自动传播修正后的注意力至整个执行过程。仿真与真实世界实验表明,GuidedAttention在位置和外观分布外(OOD)条件下均能持续提升机器人操作性能。
原文摘要 · Abstract (English)
End-to-end visuomotor policies provide little opportunity for humans to understand or correct the policy's visual attention. We propose GuidedAttention, a visuomotor imitation learning framework that introduces interpretable and correctable visual attention as an explicit intermediate representation. Task-relevant attention keypoints are predicted from camera images and condition a diffusion-based action policy. Users can inspect and optionally correct selected keypoints once at rollout initialization, after which the corrected attention is automatically propagated throughout execution by a tracking module. Experiments in simulation and the real world demonstrate that GuidedAttention consistently improves robot manipulation performance, particularly under positional and appearance out-of-distribution (OOD) conditions. https://mmurooka.github.io/guided-attention-project-page
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。