研究机器人指令中的无关信息如何影响模型表现,发现相似语义干扰最致命。
Bring the Apple, Not the Sofa: Impact of Irrelevant Context in Embodied AI Commands on VLA Models

- 通过两类语言噪声测试模型鲁棒性:改写指令与添加无关上下文。
- 相同长度下,语义相近的无关内容使性能下降约50%,远超随机内容10%的降幅。
- 提出基于大模型的过滤框架,可恢复近98.5%原始性能,适合实际部署场景。
视觉语言动作(VLA)模型广泛应用于具身智能,使机器人能够理解并执行语言指令。然而,其在真实场景中面对自然语言变化的鲁棒性尚未得到充分研究。本文系统评估了前沿VLA模型在语言扰动下的表现,包括人类生成的指令改写和添加无关上下文两种噪声类型。我们进一步根据上下文长度及其与机器人指令的语义、词汇接近度,将无关内容分为两类。研究发现,随着上下文规模扩大,模型性能持续下降;而相同长度下,语义和词汇上与指令相似的无关内容导致性能下降约50%,远高于随机上下文仅10%的降幅。人类改写指令则引发约20%的性能损失。为此,我们提出一种基于大语言模型的过滤框架,从噪声输入中提取核心命令。引入该过滤步骤后,模型在噪声条件下性能可恢复至原始水平的98.5%。
原文摘要 · Abstract (English)
Vision Language Action (VLA) models are widely used in Embodied AI, enabling robots to interpret and execute language instructions. However, their robustness to natural language variability in real-world scenarios has not been thoroughly investigated. In this work, we present a novel systematic study of the robustness of state-of-the-art VLA models under linguistic perturbations. Specifically, we evaluate model performance under two types of instruction noise: (1) human-generated paraphrasing and (2) the addition of irrelevant context. We further categorize irrelevant contexts into two groups according to their length and their semantic and lexical proximity to robot commands. In this study, we observe consistent performance degradation as context size expands. We also demonstrate that the model can exhibit relative robustness to random context, with a performance drop within 10%, while semantically and lexically similar context of the same length can trigger a quality decline of around 50%. Human paraphrases of instructions lead to a drop of nearly 20%. To mitigate this, we propose an LLM-based filtering framework that extracts core commands from noisy inputs. Incorporating our filtering step allows models to recover up to 98.5% of their original performance under noisy conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。