arXiv:2607.13429cs.ROcs.CV2026-07被引 1

通过锚定与对齐提升机器人视觉语言动作模型的泛化能力

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment

论文配图:Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment
图 1 · 摘自论文原文
  • 用冻结模型蒸馏视觉语言表示,防止微调时表征漂移
  • 将动作转为方向标签,联合训练语言与动作预测
  • 在真实机械臂上成功率提升至54%和60%,适用于长序列控制

通过行为克隆(BC)在机器人示范数据上微调预训练视觉语言模型(VLM)已成为视觉语言动作(VLA)策略的标准方法。然而,这种微调会逐步覆盖支持视觉与语义泛化的预训练表征。尽管联合训练网络图像-文本数据是常见缓解手段,但其对不同观测分别施加语言与动作损失,导致语言-动作错位,标准操控基准无法暴露此问题。本文提出Anchor-Align,通过两个新目标增强BC:视觉-语言锚定从冻结的VLM副本中逐层蒸馏表征,防止表征漂移;语言-动作对齐将每个动作目标转换为离散运动方向标签,并在同一批机器人观测上联合训练语言与动作预测。在真实xArm7机械臂上,两种主流VLA架构的实机成功率分别从28%和37%提升至54%和60%。在仿真中,对LIBERO-PRO、LIBERO-Plus和CALVIN三个数据集均显示在分布外扰动、感知鲁棒性和长时序控制上的持续改进,表明保留预训练表征与有效动作学习并非相互排斥。项目主页:anchoralignvla.github.io

原文摘要 · Abstract (English)

Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has become the standard recipe for vision-language-action (VLA) policies. However, BC finetuning progressively overwrites the pretrained representations that support visual and semantic generalization. Co-training on web image-text data, a common remedy, does not prevent this; it applies language and action losses to separate observations, leaving VLAs with language-action misalignment that standard manipulation benchmarks do not expose. We propose Anchor-Align, which augments BC with two objectives: Vision-Language Anchoring distills layer-wise representations from a frozen VLM copy to prevent this drift, while Language-Action Alignment converts each action target into a discrete motion-direction label and jointly trains language and action prediction on the same robot observation. On a physical xArm7 robot, across two widely used VLA architectures, Anchor-Align improves real-robot success on both (28% to 54% and 37% to 60%). At scale in simulation, we demonstrate consistent improvements on OOD perturbations, perceptual robustness, and long-horizon control across LIBERO-PRO, LIBERO-Plus, and CALVIN, respectively, suggesting that preserving pretrained representations and effective action learning are not fundamentally at odds. Project page: anchoralignvla.github.io

视觉语言动作机器人学习表征保持动作对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。