让视觉策略学会忽略无关信息,专注关键视觉线索。
Attentive Feature Aggregation or: How Policies Learn to Stop Worrying about Robustness and Attend to Task-Relevant Visual Cues
- 提出可学习的注意力特征聚合机制,自动筛选任务相关视觉信息
- 在仿真与真实世界中均显著提升对抗视觉干扰的鲁棒性
- 无需数据增强或微调预训练模型,轻量高效适合实际部署
采用预训练视觉表征(PVRs)已成为训练视觉-运动策略的主流方法。然而,这些强大的表示会编码大量与任务无关的场景信息,导致训练出的策略对域外视觉变化和干扰物敏感。本文将视觉-运动策略的特征池化作为解决此问题的突破口,提出轻量级、可训练的注意力特征聚合(AFA)机制,使策略能自然关注任务相关的视觉线索,忽略即使语义丰富的场景干扰。通过大量仿真与真实世界实验验证,使用AFA训练的策略在存在视觉扰动时显著优于标准池化方法,且无需昂贵的数据增强或PVR微调。结果表明,主动忽略冗余视觉信息是实现鲁棒且泛化的视觉-运动策略的关键一步。
原文摘要 · Abstract (English)
The adoption of pre-trained visual representations (PVRs), leveraging features from large-scale vision models, has become a popular paradigm for training visuomotor policies. However, these powerful representations can encode a broad range of task-irrelevant scene information, making the resulting trained policies vulnerable to out-of-domain visual changes and distractors. In this work we address visuomotor policy feature pooling as a solution to the observed lack of robustness in perturbed scenes. We achieve this via Attentive Feature Aggregation (AFA), a lightweight, trainable pooling mechanism that learns to naturally attend to task-relevant visual cues, ignoring even semantically rich scene distractors. Through extensive experiments in both simulation and the real world, we demonstrate that policies trained with AFA significantly outperform standard pooling approaches in the presence of visual perturbations, without requiring expensive dataset augmentation or fine-tuning of the PVR. Our findings show that ignoring extraneous visual information is a crucial step towards deploying robust and generalisable visuomotor policies. Project Page: tsagkas.github.io/afa
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。