让机器人只关注任务相关视觉信息,提升效率与精度
Selective Perception for Robot: Task-Aware Attention in Multimodal VLA
- 根据任务动态选择摄像头输入,仅处理有用信息
- 推理速度提升40%,控制成功率提高12.7%
- 适合资源受限的实时机器人系统使用
在机器人领域,融合多视角视觉、语言和动作信号的视觉-语言-动作(VLA)模型已成为有效方法。然而,现有方法多采用静态融合,对所有视觉输入统一处理,导致计算开销大,并引入无关背景噪声。受人类主动感知启发,本文提出一种动态信息融合框架,通过轻量级自适应路由结构,实时分析文本指令和腕部摄像头观测,预测各视角的任务相关性。对低信息效用视角,动态削减计算并仅传递必要视觉特征至策略网络,实现计算量与任务相关性的正比关系。为高效获取大规模标注数据以训练路由模块,我们构建了基于视觉-语言模型(VLMs)的自动化标注流水线,显著降低数据收集与标注成本。真实世界机器人操作实验表明,相比现有VLA模型,该方法在推理效率与控制性能上均有显著提升,验证了动态信息融合在资源受限、实时机器人控制环境中的有效性与实用性。
原文摘要 · Abstract (English)
In robotics, Vision-Language-Action (VLA) models that integrate diverse multimodal signals from multi-view inputs have emerged as an effective approach. However, most prior work adopts static fusion that processes all visual inputs uniformly, which incurs unnecessary computational overhead and allows task-irrelevant background information to act as noise. Inspired by the principles of human active perception, we propose a dynamic information fusion framework designed to maximize the efficiency and robustness of VLA models. Our approach introduces a lightweight adaptive routing architecture that analyzes the current text prompt and observations from a wrist-mounted camera in real-time to predict the task-relevance of multiple camera views. By conditionally attenuating computations for views with low informational utility and selectively providing only essential visual features to the policy network, Our framework achieves computation efficiency proportional to task relevance. Furthermore, to efficiently secure large-scale annotation data for router training, we established an automated labeling pipeline utilizing Vision-Language Models (VLMs) to minimize data collection and annotation costs. Experimental results in real-world robotic manipulation scenarios demonstrate that the proposed approach achieves significant improvements in both inference efficiency and control performance compared to existing VLA models, validating the effectiveness and practicality of dynamic information fusion in resource-constrained, real-time robot control environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。