arXiv:2506.09930cs.ROcs.CV2025-06被引 24

揭示视觉语言动作模型从意图到执行的泛化瓶颈

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models

  • 构建50个仿真任务的统一评测套件,覆盖语言、视觉与物体
  • 发现模型有强理解能力但执行精度低,尤其在分布外场景
  • 适合关注机器人泛化能力与感知-执行鸿沟的研究者

视觉语言动作(VLA)模型有望超越传统模仿学习,借助大规模视觉语言模型(VLM)的广泛泛化能力,生成多功能的“通用型”机器人策略。然而,当前对VLA的评估仍不充分:传统模仿学习基准缺乏语言指令,新兴基准任务有限且未探究预训练VLM对下游机器人策略泛化的真实贡献。此外,各机构独立搭建的实机实验阻碍了可复现性与可访问性。为此,我们提出一个包含50个仿真任务的统一探测套件,覆盖10个子类别,涵盖语言指令、视觉和物体。我们系统评估了几种主流VLA架构的泛化能力。结果表明,尽管VLM骨干网络赋予模型稳健的感知理解与高层规划能力(即良好意图),但这种能力无法可靠转化为精确的动作执行——面对分布外观测时,策略虽保持合理意图,却在动作执行上表现不佳。此外,对动作数据进行微调会削弱原有VLM的通用推理能力。我们开源任务套件与评估代码,以建立未来VLA研究的标准基准,并推动缩小感知到动作的差距。更多信息及源码见 https://ai4ce.github.io/INT-ACT/

原文摘要 · Abstract (English)

One promise that Vision-Language-Action (VLA) models hold over traditional imitation learning for robotics is to leverage the broad generalization capabilities of large Vision-Language Models (VLMs) to produce versatile, "generalist" robot policies. However, current evaluations of VLAs remain insufficient. Traditional imitation learning benchmarks are unsuitable due to the lack of language instructions. Emerging benchmarks for VLAs that incorporate language often come with limited evaluation tasks and do not intend to investigate how much VLM pretraining truly contributes to the generalization capabilities of the downstream robotic policy. Meanwhile, much research relies on real-world robot setups designed in isolation by different institutions, which creates a barrier for reproducibility and accessibility. To address this gap, we introduce a unified probing suite of 50 simulation-based tasks across 10 subcategories spanning language instruction, vision, and objects. We systematically evaluate several state-of-the-art VLA architectures on this suite to understand their generalization capability. Our results show that while VLM backbones endow VLAs with robust perceptual understanding and high level planning, which we refer to as good intentions, this does not reliably translate into precise motor execution: when faced with out-of-distribution observations, policies often exhibit coherent intentions, but falter in action execution. Moreover, finetuning on action data can erode the original VLM's generalist reasoning abilities. We release our task suite and evaluation code to serve as a standardized benchmark for future VLAs and to drive research on closing the perception-to-action gap. More information, including the source code, can be found at https://ai4ce.github.io/INT-ACT/

视觉语言动作机器人泛化感知执行鸿沟仿真评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。