用人类注视引导机器人视觉,提升效率与抗干扰能力
Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers
- 引入人类注视数据,通过聚焦关键区域降低视觉计算量
- 在高精度任务中成功率提升,计算量减少40%以上
- 适合需要高效、鲁棒视觉的机器人系统研发者
人类视觉依赖注视主动聚焦,通过中央凹机制大幅降低处理负担。相比之下,机器人系统通常对原始图像进行被动均匀处理。本文提出GIAVA(Gaze Integrated Active-Vision ALOHA),模拟人头部运动与注视调整,实现基于注视的视觉聚焦。扩展了AV-ALOHA平台,同步采集操作者的眼动、视角控制与机械臂操作数据,并开源仿真基准与数据集。受视觉变换器(ViTs)与中央凹图像分割启发,采用注视引导的分块令牌化方案,在保持性能前提下显著减少令牌数量与计算量。实验表明,该方法显著降低计算开销,增强对背景干扰的鲁棒性;在部分高精度任务中,成功率达87.6%,优于传统均匀处理。结果表明,仿人中央凹视觉具有巨大潜力,应作为机器人视觉的重要先验知识。
原文摘要 · Abstract (English)
Human vision is a highly active process driven by gaze, which directs attention to task-relevant regions through foveation, dramatically reducing visual processing. In contrast, robot learning systems typically rely on passive, uniform processing of raw camera images. In this work, we explore how incorporating human-like active gaze into robotic policies can enhance efficiency and robustness. We develop GIAVA (Gaze Integrated Active-Vision ALOHA), a robot vision system that emulates human head and neck movement, and gaze adjustment for foveated processing. Extending the AV-ALOHA robot platform, we introduce a framework for simultaneously collecting eye-tracking, perspective control, and robot manipulation demonstration data from a human operator. We also open-source a simulation benchmark and dataset for training robot policies that incorporate human gaze. Inspired by recent work in foveated image segmentation and given the widespread use of Vision Transformers (ViTs) in robot learning, we integrate gaze information into ViTs using a foveated patch tokenization scheme. Compared to uniform patch tokenization, this significantly reduces the number of tokens, and thus computation. Our results show that our method for foveated robot vision drastically reduces computational overhead, and enhances robustness to background distractors. Notably, on certain high-precision tasks, foveated vision also improves performance, as reflected in higher success rates. Together, these findings suggest that human-inspired foveated visual processing offers untapped potential and should be further considered as a useful inductive bias in robotic vision systems. https://soltanilara.github.io/giava/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。