arXiv:2606.02277cs.RO2026-06

测试视觉语言机器人模型能否真懂指令并选对物体

RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models

论文配图:RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models
图 1 · 摘自论文原文
  • 用数学和常识题设计机器人选择任务,检验理解力
  • 多数模型选对答案率接近随机,说明理解没落地到动作
  • 适合评估机器人是否真懂指令,而非靠表面线索

视觉-语言-动作(VLA)模型基于预训练语言或视觉-语言主干的语义理解能力来指导机器人动作预测。然而,机器人微调通常以特定任务的动作分布为模仿目标,许多评估可被视觉或指令-动作捷径解决。我们提出RoboSemanticBench(RSB),一个用于诊断动作预测中语义接地的具身基准:后训练的VLA模型能否利用复杂指令语义选择并操作正确的物理目标。每个回合中,机器人接收多选数学题或常识题,观察候选答案块,需抓取对应正确答案的块。RSB涵盖控制性算术、小学数学理解以及四选一和十选一情境下的常识或事实理解。在代表性VLA模型上,我们发现尽管许多策略能成功抓取候选块,但在控制抓取成功率后,其选择语义正确块的准确率接近随机甚至更低,揭示了主干级语义能力与动作预测之间持续存在的差距。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models are built on the premise that semantic understanding from pretrained language or vision-language backbones should guide robot action prediction. Yet robot fine-tuning is optimized as imitation over task-specific action distributions, and many evaluations can be solved through visual or instruction-action shortcuts. We introduce RoboSemanticBench (RSB), an embodied benchmark for diagnosing semantic grounding in action prediction: whether post-trained VLA models can use complex instruction semantics to select and manipulate the correct physical target. In each episode, a robot receives a multiple-choice math or general-knowledge question, observes candidate answer blocks, and must grasp the block corresponding to the correct answer. RSB covers controlled arithmetic, grade-school mathematical understanding, and commonsense or factual understanding under four-choice and ten-choice suites. Across representative VLA models, we find that many policies learn to grasp candidate blocks but select the semantically correct block at near-random or below-random rates after controlling for grasp success, revealing a persistent gap between backbone-level semantic competence and action prediction.

机器人语义理解视觉语言基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。