提出新基准LangGap,揭示视觉语言动作模型严重忽视语言指令
LangGap: Diagnosing and Closing the Language Gap in Vision-Language-Action Models
- 设计四维语义扰动方法,固定场景变动指令测试模型理解力
- 现有模型在新任务上成功率从0%提升至90%,仍难应对复杂语言变化
- 适合研究多模态理解、机器人指令泛化能力的学者参考
视觉-语言-动作(VLA)模型在标准基准上成功率超95%。但系统实验发现,当前最先进模型严重忽略语言指令。既有工作缺乏:(1) 系统性语义扰动诊断,(2) 强制语言理解的基准设计,(3) 语言多样性训练数据。本文构建LangGap基准,采用四维语义扰动方法——固定桌面上下文,改变指令语义——揭示π0.5模型的语言理解缺陷。现有基准如LIBERO每场景仅设一个任务,未充分利用物体与目标位置;LangGap在相同布局下充分多样化拾取放置任务,迫使模型真正理解语言。实验表明,针对性数据增强可部分弥补语言差距:单任务训练成功率由0%升至90%,多任务训练由0%升至28%。但随着扩展任务语义多样性增加,模型学习能力严重不足,甚至已训练任务表现也差。这揭示了VLA模型在理解多样化语言指令上的根本挑战,正是LangGap长期价值所在。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models achieve over 95% success on standard benchmarks. However, through systematic experiments, we find that current state-of-the-art VLA models largely ignore language instructions. Prior work lacks: (1) systematic semantic perturbation diagnostics, (2) a benchmark that forces language understanding by design, and (3) linguistically diverse training data. This paper constructs the LangGap benchmark, based on a four-dimensional semantic perturbation method -- varying instruction semantics while keeping the tabletop layout fixed -- revealing language understanding deficits in π0.5. Existing benchmarks like LIBERO assign only one task per layout, underutilizing available objects and target locations; LangGap fully diversifies pick-and-place tasks under identical layouts, forcing models to truly understand language. Experiments show that targeted data augmentation can partially close the language gap -- success rate improves from 0% to 90% with single-task training, and 0% to 28% with multi-task training. However, as semantic diversity of extended tasks increases, model learning capacity proves severely insufficient; even trained tasks perform poorly. This reveals a fundamental challenge for VLA models in understanding diverse language instructions -- precisely the long-term value of LangGap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。