arXiv:2505.05540cs.CVcs.LG2025-05被引 6

评测视觉语言动作模型在生成环境中的泛化能力,发现其零样本表现仍有瓶颈。

Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments

  • 构建多任务生成环境基准,评估模型在分布外任务上的表现。
  • 所有模型在分布外任务上表现显著下降,性能受动作表示与任务复杂度影响。
  • 语言动作模型优于纯视觉模型,提示工程对性能影响极大。

视觉-语言-动作(VLA)模型通过融合视觉感知、语言理解和动作执行,推动通用机器人系统的发展。然而,对这些模型的系统性评估,尤其是其在程序生成的分布外(OOD)环境中的零样本泛化能力,仍显不足。本文提出MultiNet v0.2,一个综合性基准,用于评估最先进的视觉语言模型(VLM)和视觉语言动作模型(VLA)——包括GPT-4o、GPT-4.1、OpenVLA、Pi0 Base和Pi0 FAST——在Procgen基准中多样化程序化任务上的泛化性能。分析显示:(1)所有被测模型在零样本泛化到分布外任务时均存在显著局限,性能受动作表示方式和任务复杂度显著影响;(2)由于架构设计更稳健,VLA模型总体表现优于其他模型;(3)在合理约束下,VLM变体性能显著提升,凸显模型表现对精确提示工程的高度敏感性。我们公开发布该基准、评估框架及研究发现,以支持未来VLA模型的评估,并识别其在分布外数字任务应用中的关键改进方向。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models represent an important step toward general-purpose robotic systems by integrating visual perception, language understanding, and action execution. However, systematic evaluation of these models, particularly their zero-shot generalization capabilities in procedurally out-of-distribution (OOD) environments, remains limited. In this paper, we introduce MultiNet v0.2, a comprehensive benchmark designed to evaluate and analyze the generalization performance of state-of-the-art VLMs and VLAs - including GPT-4o, GPT-4.1, OpenVLA, Pi0 Base, and Pi0 FAST - on diverse procedural tasks from the Procgen benchmark. Our analysis reveals several critical insights: (1) all evaluated models exhibit significant limitations in zero-shot generalization to OOD tasks, with performance heavily influenced by factors such as action representation and task complexity; (2) VLAs generally outperforms other models due to their robust architectural design; and (3) VLM variants demonstrate substantial improvements when constrained appropriately, highlighting the sensitivity of model performance to precise prompt engineering. We release our benchmark, evaluation framework, and findings to enable the assessment of future VLA models and identify critical areas for improvement in their application to out-of-distribution digital tasks.

视觉语言动作泛化能力基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。