arXiv:2603.12772cs.CVcs.LG2026-03被引 2

让视觉语言动作模型更好理解时间信息,提升复杂操作能力。

PVI: Plug-in Visual Injection for Vision-Language-Action Models

  • 轻量级模块通过残差路径注入视觉特征,不改动原有结构
  • 使用视频特征比图像特征提升更大,尤其在多阶段任务中
  • 适合需要状态追踪的复杂机器人操作,实机验证有效

将预训练视觉语言模型(VLM)与流匹配动作专家结合的视觉-语言-动作(VLA)架构已成为语言控制操作的强大范式。然而,为语义抽象优化的VLM通常削弱细粒度几何线索,并缺乏动作专家所需的显式时间证据。现有方法虽尝试注入辅助视觉特征,但多集中于静态空间表示,或需大幅修改架构以处理时序输入,导致时间信息未被充分探索。我们提出插件式视觉注入(PVI),一种轻量级、编码器无关的模块,可附加至预训练动作专家,通过零初始化残差路径注入辅助视觉表示,仅需单阶段微调即可保留预训练行为。实验表明,PVI在基础策略和多种替代注入方法上均取得稳定提升;受控研究显示,时序视频特征(V-JEPA2)优于强静态图像特征(DINOv2),且在需状态追踪与协调的多阶段任务中收益最大。真实机器人长时程双臂折叠布料实验进一步验证了PVI在真实场景中的实用性。

原文摘要 · Abstract (English)

VLA architectures that pair a pretrained VLM with a flow-matching action expert have emerged as a strong paradigm for language-conditioned manipulation. Yet the VLM, optimized for semantic abstraction and typically conditioned on static visual observations, tends to attenuate fine-grained geometric cues and often lacks explicit temporal evidence for the action expert. Prior work mitigates this by injecting auxiliary visual features, but existing approaches either focus on static spatial representations or require substantial architectural modifications to accommodate temporal inputs, leaving temporal information underexplored. We propose Plug-in Visual Injection (PVI), a lightweight, encoder-agnostic module that attaches to a pretrained action expert and injects auxiliary visual representations via zero-initialized residual pathways, preserving pretrained behavior with only single-stage fine-tuning. Using PVI, we obtain consistent gains over the base policy and a range of competitive alternative injection strategies, and our controlled study shows that temporal video features (V-JEPA2) outperform strong static image features (DINOv2), with the largest gains on multi-phase tasks requiring state tracking and coordination. Real-robot experiments on long-horizon bimanual cloth folding further demonstrate the practicality of PVI beyond simulation.

视觉注入动作模型时序特征机器人操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。