让视觉语言动作模型聚焦关键视觉信息,提升机器人操作精度。
FocusVLA: Focused Visual Utilization for Vision-Language-Action Models
- 用级联注意力机制消除模型对无关视觉的依赖
- 动态选择任务相关视觉区域,减少噪声干扰
- 在仿真与真实机器人任务中均显著提速增效
视觉-语言-动作(VLA)模型通过结合丰富的视觉语言信息来改进动作生成,但当前自回归策略面临三大瓶颈:架构偏见导致忽略视觉细节,过多视觉标记使注意力难以聚焦正确区域,无关视觉信息引入大量噪声,严重损害动作质量。本文实证验证这些问题,并指出性能受限于视觉信息利用方式而非表示质量。为此提出FocusVLA新范式,引导模型关注任务相关的视觉区域,有效连接视觉与动作。核心包括:模态级联注意力,消除捷径路径,迫使模型依赖任务相关的视觉细节;焦点注意力,动态选择相关视觉块,控制信息量并显式抑制无关噪声。在多种模拟与真实世界机器人基准测试中,FocusVLA不仅更高效利用视觉细节完成灵巧操作,还在多个任务上显著提升性能并加速收敛。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models improve action generation by conditioning policies on rich vision-language information. However, current auto-regressive policies are constrained by three bottlenecks: (1) architectural bias drives models to overlook visual details, (2) an excessive number of visual tokens makes attention difficult to focus on the correct regions, and (3) task-irrelevant visual information introduces substantial noise - together severely impairing the quality of action. In this paper, we investigate how to effectively utilize different visual representations for action generation. To this end, we first empirically validate the above issues and show that VLA performance is primarily limited by how visual information is utilized, rather than by the quality of visual representations. Based on these insights, we introduce FocusVLA, a novel paradigm that directs the model's attention to task-relevant visual regions to effectively bridge vision to action. Specifically, we first propose Modality Cascaded Attention to eliminate shortcut pathways, thereby compelling VLA models to rely on task-relevant visual details for action generation. Furthermore, we propose Focus Attention, which dynamically selects task-relevant visual patches to control information quantity while explicitly modulating their influence to suppress task-irrelevant noise. Extensive experiments on both simulated and real-world robotic benchmarks demonstrate that FocusVLA not only effectively leverages visual details to perform dexterous manipulations, but also substantially improves performance and accelerates convergence across a variety of tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。