提升机器人视觉-语言-动作模型的感知精度,解决视觉信息丢失问题
V-Link: Recovering Lost Visual Representations in Action DiT for Vision-Language-Action Models

- 通过空间与语义查询显式恢复视觉特征,增强动作生成的感知基础
- 在多个数据集上平均成功率提升1.9%至31.2%,真实场景任务提升20%-24%
- 适合关注机器人精细操作与多模态对齐的开发者和研究者
视觉-语言-动作(VLA)模型通过整合视觉感知、语言理解与连续动作控制,为通用机器人操作提供了可扩展路径。然而,我们发现现有VLA架构存在关键缺陷:动作专家对视觉语言模型(VLM)中的3D几何与2D语义信息访问有限,导致感知基底弱化,影响细粒度操作性能。为此,我们提出V-Link,通过在视觉-语言到动作的特征传递过程中显式恢复视觉表征。具体而言,V-Link在VLM中学习互补的空间与语义查询,并通过非对称路径注入到动作DiT中:语义查询补充原始图像令牌,空间查询则提供专用于空间定位的动作生成条件。在LIBERO、LIBERO-Plus与RoboTwin 2.0上,V-Link相较基线模型GR00T N1.6分别实现+1.9%、+31.2%与+18.8%的平均成功率提升;在真实世界的人形机器人任务AGIBOT A3 Ultra上,成功率进一步提升+20%与+24%。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models provide a scalable path toward generalist robotic manipulation by integrating visual perception, language understanding, and continuous action control. However, we reveal a critical limitation of VLA architectures: the action expert has limited access to the 3D geometric and 2D semantic information available in VLM features. This accessibility gap weakens perceptual grounding and limits performance on fine-grained robotic manipulation. To address this issue, we propose V-Link, which explicitly recovers visual representations during the vision-language (VL) to action (A) feature transfer. Specifically, V-Link learns complementary Spatial and Semantic Query representations within the VLM and injects them into Action DiT through asymmetric pathways. Semantic Queries complement the original VLM image tokens, whereas Spatial Queries provide dedicated geometric conditioning for spatially grounded action generation. Across LIBERO, LIBERO-Plus, and RoboTwin 2.0, our V-Link improves the average success rate over base model GR00T N1.6 by +1.9%, +31.2%, and +18.8%, respectively. On the AGIBOT A3 Ultra, V-Link further achieves gains of +20% and +24% on two real-world humanoid tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。