构建文本不可替代的机器人指令执行评测基准,逼真检验模型是否真听懂了指令。
InstructMove: A Text-Indispensable Benchmark for Instruction-Following Manipulation

- 设计包含语义干扰物的抓放场景,让唯一正确动作依赖语言理解
- 在多个任务上验证主流视觉语言模型仍会利用视觉捷径而非理解指令
- 支持训练-评估分离和语言依赖诊断,适合研究真实指令跟随能力
视觉-语言-动作(VLA)模型通过将机器人动作与自然语言指令关联,使通用机器人操作成为可能。但现有操作基准常无法充分测试模型是否真正遵循指令:目标物体或位置往往视觉显著或唯一可行,使模型可不依赖语言也能成功。我们主张指令跟随评估应具备文本不可替代性——多个动作在视觉和物理上均合理,仅一个符合语言描述。为此提出InstructMove基准,通过含语义干扰物的抓放场景,将指令遵循分解为类别识别、属性区分、空间推理和组合抓放。该基准支持训练-评估分离协议,使用训练数据和保留评估任务,并提供语言依赖性诊断。对代表性VLA策略的实验表明,InstructMove能有效检测视觉捷径问题,且其仿真数据可提升真实世界指令跟随性能。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models have made general-purpose robot manipulation increasingly plausible by conditioning robot actions on natural-language instructions. A key test of such generality is whether policies actually follow language instructions. Yet many manipulation benchmarks leave this ability underdetermined: the intended object or destination is often visually salient or uniquely feasible, allowing policies to succeed without grounding the instruction. We argue that instruction-following evaluation should be text-indispensable: multiple actions should be visually and physically plausible, while only one should be consistent with the language instruction. We introduce InstructMove, a text-indispensable benchmark for instruction-following manipulation. InstructMove instantiates this principle in pick-and-place scenes with semantic distractors, decomposing instruction following into category identification, attribute discrimination, spatial reasoning, and compositional pick-and-place. InstructMove supports a train-eval protocol with InstructMove training data and held-out evaluation tasks, with additional diagnostics for language dependence. Experiments with representative VLA policies show that InstructMove provides a controlled testbed for diagnosing visual shortcuts and that InstructMove simulation data can improve real-world instruction-following manipulation performance. Code: https://github.com/HorizonRobotics/RoboOrchardSim
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。