现有视觉语言动作模型的物理推理能力无法被验证,因评估指标无法区分语义理解与真实动作决策。
Position: Vision-Language-Action Models Cannot Be Verified to Perform Physical Reasoning
- 将模型行为拆解为语义映射与物理动作决策,揭示评估盲区。
- 任务成功率无法区分语义匹配、分布重叠与真正物理泛化。
- 适合关注机器人泛化评估方法论的研究者阅读。
基于预训练视觉语言模型(VLM)的视觉语言动作(VLA)系统在机器人操作基准上表现迅速提升。这些进展常被解读为互联网规模数据学习的语义表征能有效迁移至物理执行泛化。本文指出,这一解释所依赖的核心假设——语义泛化足以支持物理动作决策——尚未被独立验证,且在当前评估协议下无法测试。我们通过将VLA策略分解为语义映射与物理动作决策两部分,证明任务成功率这一主流评估指标无法区分这两种能力来源。因此,基准性能提升可能源于语义匹配、分布重叠或真正的物理泛化等多种解释。我们进一步指出,这种可识别性缺口因叙事漂移而加剧:后续系统沿用并强化了先前对性能提升的解释,却未分离出潜在因果机制。为此,我们提出一种新研究方向:通过引入可控变量的评估设计,分别测量语义与物理泛化能力。此类设计无需访问模型内部即可因果归因,并实证检验VLM作为语义接口而非隐式物理能力来源的角色。本文目标并非否定VLM在机器人中的作用,而是明确物理泛化声明可被有意义评估的前提条件。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) systems, built on pretrained vision-language models (VLMs), have shown rapidly improving performance on robot manipulation benchmarks. These gains are commonly interpreted as evidence that semantic representations learned from internet-scale data transfer to physical execution generalization. This position paper argues that the assumption underlying this interpretation -- that semantic generalization is sufficient to support physical action decisions -- has not been independently verified and cannot be tested under current evaluation protocols. We support this claim by decomposing VLA policies into semantic mapping and physical action decision, and showing that task success rate -- the dominant evaluation metric -- cannot distinguish between these two sources of capability. As a result, improvements in benchmark performance are consistent with multiple competing explanations, including semantic matching, distributional overlap, and genuine physical generalization. We further argue that this identifiability gap has been reinforced through narrative drift, whereby successive systems inherit and strengthen prior interpretations of performance gains without isolating the underlying causal mechanism. To address this limitation, we propose a research direction based on evaluation designs that introduce controlled variation to separately measure semantic and physical generalization. Such designs make it possible to causally attribute performance without requiring access to model internals, and to empirically assess the role of VLM backbones as semantic interfaces rather than implicit sources of physical competence. Our goal is not to refute the role of VLMs in robotics, but to clarify the conditions under which claims of physical generalization can be meaningfully evaluated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。