发现视觉语言动作模型在非英语指令上表现大幅下降,提出对齐方法改善多语言能力。
Beyond English: Uncovering the Multilingual Gap in Vision-Language-Action Models

- 构建多语言指令数据集,系统评估主流视觉语言动作模型的跨语言能力。
- 模型在非英语指令上性能显著下降,即使底层语言模型支持多语种。
- 提出主成分对齐方法,通过特征空间对齐有效缩小多语言性能差距。
视觉语言动作(VLA)模型近年来在从大规模多模态数据中学习通用机器人策略方面展现出潜力。然而,大多数现有VLA系统主要以英文指令进行训练和评估,其对其他语言指令的理解与执行能力仍缺乏探索。尽管底层大语言模型通常具备多语言能力,但这种能力是否能在训练过程中传递给VLA尚不明确。本文首次系统研究了VLA模型中的多语言指令遵循问题。我们通过翻译现有基准数据集的指令,构建多语言指令集,并在模拟环境中评估多个代表性VLA模型。实验揭示显著的多语言差距:仅用英文训练的模型在其他语言上的表现大幅下降,即便其语言主干支持多语种。跨语言迁移行为分析表明,性能下降同时源于指令理解与动作执行两方面。表示分析显示,多语言指令引发的表征偏移可能加剧该差距。基于此,我们提出一种简单有效的多语言微调方法——多语言主成分对齐(Multilingual Principal Component Alignment),利用主成分分析提取主成分子空间并对齐多语言表示,显著减小多语言性能差距。
原文摘要 · Abstract (English)
Vision-Language-Action models have recently demonstrated promising capabilities in learning generalist robot policies from large-scale multimodal data. However, most existing VLA systems are trained and evaluated primarily with English instructions, leaving their ability to understand and execute instructions in other languages largely unexplored. While the underlying large language models often possess multilingual capabilities, it remains unclear whether these multilingual capabilities transfer to VLAs during training. In this work, we present the first systematic study of multilingual instruction following in VLA models. We first construct multilingual instructions by extending existing benchmarks with translations of their instructions. Using these instructions, we evaluate several representative VLA models across a range of tasks in simulation settings. Our experiments reveal a significant multilingual gap: models trained primarily on English instructions exhibit substantial performance degradation when evaluated on other languages, even when the underlying language backbone is multilingual. We provide several findings and analyses to understand the multilingual gap. Cross-lingual transfer behavior analysis shows that performance drops correlate with both instruction understanding and action execution. Representation analyses suggest that multilingual instruction-caused representation shifts may contribute to the multilingual gap. Motivated by these findings, we further explore strategies to improve multilingual performance in VLAs. We propose a simple yet effective multilingual fine-tuning approach, Multilingual Principal Component Alignment, which leverages Principal Component Analysis to get the principal component subspace and align projected multilingual representations, effectively reducing the multilingual performance gap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。