arXiv:2510.25616cs.LGcs.AI2025-10被引 31

发现视觉语言模型在适配动作任务时会丢失视觉表征,提出有效恢复方法。

Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization

  • 通过分析隐藏层与注意力图,揭示动作微调导致视觉表征退化
  • 设计对比任务验证微调对视觉-语言能力的影响,量化退化程度
  • 提出简单有效的对齐方法,显著提升模型在分布外场景的泛化能力

视觉-语言-动作(VLA)模型的成功源于预训练视觉-语言模型(VLM)赋予智能体可迁移的世界知识与视觉-语言(VL)对齐能力,为具备更强泛化性的动作模型奠定基础。然而当这些VLM被适配至动作模态时,其原有的视觉-语言表征与知识保留程度仍不明确。本文系统研究了VLA微调过程中的表征保留问题,发现直接进行动作微调会导致视觉表征退化。通过探测隐藏表示并分析注意力图,我们设计了一系列针对性任务与方法,对比VLA模型与其对应VLM,隔离出动作微调引发的视觉-语言能力变化。进一步评估多种视觉表征对齐策略,提出一种简单而高效的方法,缓解退化现象,并实现更优的分布外(OOD)泛化性能。整体分析揭示了动作微调与视觉-语言表征退化之间的权衡关系,提出了实用的恢复继承视觉-语言能力的路径。代码已公开:https://blind-vla-paper.github.io

原文摘要 · Abstract (English)

The growing success of Vision-Language-Action (VLA) models stems from the promise that pretrained Vision-Language Models (VLMs) can endow agents with transferable world knowledge and vision-language (VL) grounding, laying a foundation for action models with broader generalization. Yet when these VLMs are adapted to the action modality, it remains unclear to what extent their original VL representations and knowledge are preserved. In this work, we conduct a systematic study of representation retention during VLA fine-tuning, showing that naive action fine-tuning leads to degradation of visual representations. To characterize and measure these effects, we probe VLA's hidden representations and analyze attention maps, further, we design a set of targeted tasks and methods that contrast VLA models with their counterpart VLMs, isolating changes in VL capabilities induced by action fine-tuning. We further evaluate a range of strategies for aligning visual representations and introduce a simple yet effective method that mitigates degradation and yields improved generalization to out-of-distribution (OOD) scenarios. Taken together, our analysis clarifies the trade-off between action fine-tuning and the degradation of VL representations and highlights practical approaches to recover inherited VL capabilities. Code is publicly available: https://blind-vla-paper.github.io

视觉-语言模型微调泛化性对齐方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。