arXiv:2606.11906cs.CL2026-06ACL被引 1

多语言测试发现视觉语言动作模型对指令敏感度不均,影响机器人任务成功率。

When Does Language Matter? Multilingual Instructions Reveal Step-wise Language Sensitivity in Vision-Language-Action Models

论文配图:When Does Language Matter? Multilingual Instructions Reveal Step-wise Language Sensitivity in Vision-Language-Action Models
图 1 · 摘自论文原文
  • 通过十种语言翻译基准测试,分析模型在不同步骤中的语言依赖性。
  • 非英语指令下成功率下降30-50%,关键步骤语言敏感导致整体失败。
  • 提出按步骤动态调整表示的方法,提升模型在多语言下的鲁棒性。

视觉语言动作(VLA)模型在语言引导的机器人操作中表现优异,但其对语言变化的鲁棒性尚不明确。本文首次系统性地将LIBERO基准翻译为十种语言进行多语言评估,发现非英语指令下性能显著下降,成功率达30-50%。细粒度分析显示,语言影响在任务步骤间分布极不均匀:部分步骤高度依赖语言,是任务失败的主要原因;而其他步骤则基本与语言无关。基于此,我们提出一种分步推理时干预策略,根据各步骤的语言敏感度对表征进行对齐,显著提升模型在语言变异下的表现。结果表明,VLA模型的语言鲁棒性本质上是分步控制问题,强调了时序结构分析对可靠具身智能体的重要性。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have shown strong performance in language-conditioned robotic manipulation, yet their robustness to linguistic variation remains poorly understood. In this work, we present the first systematic multilingual evaluation of VLA models by translating the LIBERO benchmark into ten languages, revealing severe performance degradation under non-English instructions, with success rates dropping by 30-50%. Through fine-grained analysis of task executions, we find that language influence is highly non-uniform across steps: certain steps exhibit strong language dependence and dominate overall task failure, while others are largely language-agnostic. Based on this insight, we propose a step-wise inference-time intervention that aligns representations according to step language sensitivity, substantially improving performance under linguistic variation. Our results indicate that language robustness in VLA models is fundamentally a step-wise control problem, highlighting the importance of temporally structured analysis for reliable embodied agents.

多语言机器人视觉语言动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。