让机器人自动找关键视觉区域,提升操控的效率与鲁棒性。
Attention from Action, for Action: Emergent Visual Bottlenecks for Policy Learning

- 基于动作反馈学习动态关注区域,无需人工标注位置。
- 实机测试中任务成功率从48.3%提升至76.7%,光照变化下从20.0%升至60.0%。
- 适用于对数据效率和环境适应性要求高的真实机器人控制场景。
通过聚焦视觉输入到感兴趣区域(ROIs),视觉瓶颈可提升视觉-运动学习的数据效率,将‘看哪里’与‘如何行动’解耦。现有方法多依赖外部空间标签(如视线、物体类别或功能标注)。无标签方案通常基于轨迹检测夹爪或运动事件,以末端执行器投影点为中心生成固定裁剪区域。此类动作衍生裁剪虽无需额外标注,但其事件时机、代理点和裁剪尺度为固定选择,当控制所需视觉信息偏离末端执行器或随任务进程持续变化时,易导致裁剪错位。本文提出Seeker:一种任务与状态条件化的读出机制,从动作监督中学习注意力。基于冻结的DINOv3特征,Seeker 迭代更新查询,融合视觉证据,仅凭动作信号生成进度感知的ROIs。该学习到的ROI可用于RGB裁剪、掩码引导的背景增强及点云过滤。在仿真与真实世界中,Seeker均优于无裁剪、增强和动作衍生裁剪基线。在真实机器人上,其平均域内成功率从最优基线的48.3%提升至76.7%,在光照/背景变化下成功率从20.0%提升至60.0%。
原文摘要 · Abstract (English)
Visual bottlenecks that focus policy inputs on regions of interest (ROIs) can improve data-efficient visuomotor learning by separating where to look from how to act. Many ROI interfaces rely on external spatial labels, such as gaze, object classes, or affordance annotations. Label-free alternatives often derive crops from trajectories by detecting gripper or motion events and centering a fixed crop at the projected end-effector. Such action-derived crops are useful spatial priors that require no additional labels, but they encode fixed choices about event timing, proxy points, and crop scale. When the visual evidence needed for control lies away from the end-effector or changes continuously with task progress, these crops can become misaligned. We propose Seeker, a task- and state-conditioned readout that learns attention from action. Starting from frozen DINOv3 features, Seeker iteratively updates a query with gathered visual evidence, producing progression-aware ROIs solely from action supervision. The learned ROI serves as a spatial interface for RGB cropping, mask-guided background augmentation, and point-cloud filtering. In simulation and the real world, Seeker improves data efficiency and robustness over no-crop, augmentation, and action-derived crop baselines. On real robots, Seeker raises average in-domain success from the best baseline's 48.3% to 76.7% and success under lighting/background shifts from 20.0% to 60.0%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。