arXiv:2604.18000cs.RO2026-04

提出新基准BeTTER,揭露视觉语言动作模型在真实物理任务中的推理幻觉。

Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models

  • 设计因果干预测试,隔离执行误差,精准诊断高层推理缺陷。
  • 顶尖模型在动态场景中失败率超80%,暴露出语义坍塌与行为惯性。
  • 揭示架构瓶颈:容量压缩和短视下采样导致表征退化,适合关注机器人真实智能的研究者。

近期视觉语言动作(VLA)模型在标准机器人基准上表现优异,引发对通用物理智能的乐观预期。然而,现有证据表明,基准成功与真实具身推理之间存在系统性偏差,高分是否反映真正认知能力存疑。为此,我们提出BeTTER——一个用于测试机器人策略中真正具身推理能力的诊断基准。BeTTER通过施加定向因果干预(如空间布局变化、时间外推),并强制运动学隔离,明确分离高层推理失败与底层执行限制。系统评估显示,当前最先进VLAs在动态场景中表现灾难性,出现严重词汇-运动捷径、行为惯性及语义特征坍塌。关键机制分析表明,这些现象源于根本性架构瓶颈,如容量压缩和短视下采样,系统性削弱模型的基础语义表征。我们证明,高度静态的评估协议通过允许模型过拟合感知运动先验,有效掩盖了这一退化。结合真实机器人验证,研究确认该表征崩溃并非仿真伪影,凸显未来VLA范式亟需解决高频控制与高层推理之间的结构性矛盾。

原文摘要 · Abstract (English)

Recent Vision-Language-Action (VLA) models report impressive success rates on standard robotic benchmarks, fueling optimism about general-purpose physical intelligence. However, recent evidence suggests a systematic misalignment between standard benchmark success and true embodied reasoning, raising the question of whether these high scores reflect genuine cognitive capability. To address this gap, we introduce BeTTER, a diagnostic Benchmark for Testing True Embodied Reasoning in robotic policies. BeTTER applies targeted causal interventions (e.g., spatial layout shifts, temporal extrapolation) while enforcing kinematic isolation to explicitly decouple high-level reasoning failures from low-level execution limits. Through systematic evaluation, we reveal that state-of-the-art VLAs catastrophically fail in dynamic scenarios, exhibiting severe lexical-kinematic shortcuts, behavioral inertia, and semantic feature collapse. Crucially, our mechanistic analysis traces these symptoms to fundamental architectural bottlenecks - such as capacity compression and myopic downsampling - which systematically degrade the model's foundational semantic representation. We demonstrate that highly static evaluation protocols effectively mask this degradation by allowing optimization to overfit to sensorimotor priors. Supported by real-world robotic validation, our findings confirm that this representational breakdown is not a simulation artifact, highlighting the critical need for future VLA paradigms to resolve the structural tension between high-frequency control and high-level reasoning.

机器人具身智能评估基准推理幻觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。