arXiv:2603.15618cs.CV2026-03被引 4

提升视觉语言动作模型的视觉理解能力,让机器人更精准执行复杂操作。

Look Before Acting: Enhancing Vision Foundation Representations for Vision-Language-Action Models

  • 通过共享注意力机制,将多层级视觉特征注入动作模型深层,增强视觉表征。
  • 在模拟和真实任务上分别提升9.0%和7.5%的性能,显著优于现有方法。
  • 适合关注机器人视觉决策、多模态模型优化的研究者与开发者。

视觉-语言-动作(VLA)模型已成为机器人操作的新兴范式,其可靠的动作预测高度依赖于对视觉观察与语言指令的准确理解与融合。尽管已有研究致力于提升VLA模型的视觉能力,但多数方法将大语言模型(LLM)视为黑箱,难以揭示视觉信息如何被嵌入动作生成。为此,我们系统分析了多种VLA模型在不同动作生成范式下的表现,发现视觉标记在深层网络中的敏感度逐步下降。基于此,我们提出DeepVision-VLA,采用视觉-语言混合变换器(VL-MoT)框架,实现视觉基础模型与VLA主干之间的共享注意力,将视觉专家的多层级特征注入到VLA主干深层,以增强视觉表示,支持精确复杂的操作。此外,我们引入动作引导视觉剪枝(AGVP),利用浅层注意力剔除无关视觉标记,保留任务相关关键线索,计算开销极小。DeepVision-VLA在模拟和真实世界任务中分别领先于当前最优方法9.0%和7.5%,为视觉增强型VLA模型的设计提供了新思路。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for robotic manipulation, in which reliable action prediction critically depends on accurately interpreting and integrating visual observations conditioned on language instructions. Although recent works have sought to enhance the visual capabilities of VLA models, most approaches treat the LLM backbone as a black box, providing limited insight into how visual information is grounded into action generation. Therefore, we perform a systematic analysis of multiple VLA models across different action-generation paradigms and observe that sensitivity to visual tokens progressively decreases in deeper layers during action generation. Motivated by this observation, we propose \textbf{DeepVision-VLA}, built on a \textbf{Vision-Language Mixture-of-Transformers (VL-MoT)} framework. This framework enables shared attention between the vision foundation model and the VLA backbone, injecting multi-level visual features from the vision expert into deeper layers of the VLA backbone to enhance visual representations for precise and complex manipulation. In addition, we introduce \textbf{Action-Guided Visual Pruning (AGVP)}, which leverages shallow-layer attention to prune irrelevant visual tokens while preserving task-relevant ones, reinforcing critical visual cues for manipulation with minimal computational overhead. DeepVision-VLA outperforms prior state-of-the-art methods by 9.0\% and 7.5\% on simulated and real-world tasks, respectively, providing new insights for the design of visually enhanced VLA models.

视觉语言动作机器人操作多模态模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。