用立体视觉注意力提升机器人在变化视角下的实时操作鲁棒性
Stereo Multistage Spatial Attention for Real-Time Mobile Manipulation Under Visual Scale Variation and Disturbances

- 通过多阶段立体空间注意力提取关键视觉特征
- 在四种真实任务中成功率显著优于基线方法
- 适合需要高鲁棒性的移动操作场景研究者
在开放、非结构化的真实环境中,机器人需依赖机载视觉感知并自主移动。持续变化的相机视角导致目标物体视觉尺度剧烈波动,影响基于视觉的动作生成。本文提出一种基于立体多阶段空间注意力的深度预测学习方法,用于实时移动操作。该方法从立体图像中提取任务相关空间注意力点,并通过分层循环架构与机器人状态融合,实现闭环动作预测。我们在一台移动机械臂上评估了四个真实世界任务,包括刚性放置、可动物体操作和柔性物体交互。在随机初始位置和视觉干扰条件下实验表明,相比代表性模仿学习与视觉-语言-动作基线,在相同控制设置下,本方法展现出更高的鲁棒性和任务成功率。结果表明,结构化的立体空间注意力结合预测时序建模,是所评估移动操作场景中的有效解决方案。
原文摘要 · Abstract (English)
Robots operating in open, unstructured real-world environments must rely on onboard visual perception while autonomously moving across different locations. Continuous changes in onboard camera viewpoints cause significant visual scale variations in target objects, affecting vision-based motion generation. In this work, we present a stereo multistage spatial attention-based deep predictive learning method for real-time mobile manipulation. The proposed methods extracts task-relevant spatial attention points from stereo images and integrates them with robot states through a hierarchical recurrent architecture for closed-loop action prediction. We evaluate the system on four real-world mobile manipulation tasks using a mobile manipulator, including rigid placement, articulated object manipulation, and deformable object interaction. Experiments under randomized initial positions and visual disturbance conditions demonstrate improved robustness and task success rates compared to representative imitation learning and vision-language-action baselines under identical control settings. The results indicate that structured stereo spatial attention combined with predictive temporal modeling provides an effective solution within the evaluated mobile manipulation scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。