评测三大视觉语言动作模型在20个机器人任务中的表现,发现性能差异大且难处理复杂操作。
Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks
- 构建跨任务、跨平台的评估框架,测试3个先进模型在20个数据集上的表现。
- GPT-4o通过提示工程实现最稳定表现,但所有模型均难完成多步规划任务。
- 模型性能受动作空间和环境因素影响显著,适合研究通用机器人系统者参考。
视觉-语言-动作(VLA)模型为构建通用机器人系统提供了有前景的方向,具备融合视觉理解、语言理解和动作生成的能力。然而,对这些模型在多样化机器人任务中的系统性评估仍显不足。本文提出一个全面的评估框架与基准套件,对三种最先进的视觉语言模型(VLM)和VLA——GPT-4o、OpenVLA和JAT——在来自Open-X-Embodiment数据集集合的20个不同数据集上进行评估,考察其在多种操作任务中的表现。分析揭示若干关键发现:1. 当前VLA模型在不同任务和机器人平台上表现出显著性能差异,其中GPT-4o通过高级提示工程展现出最一致的表现;2. 所有模型在需要多步规划的复杂操作任务中均表现不佳;3. 模型性能明显受动作空间特性与环境因素影响。我们公开评估框架与结果,以促进未来VLA模型的系统性评估,并识别通用机器人系统发展的关键改进方向。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models represent a promising direction for developing general-purpose robotic systems, demonstrating the ability to combine visual understanding, language comprehension, and action generation. However, systematic evaluation of these models across diverse robotic tasks remains limited. In this work, we present a comprehensive evaluation framework and benchmark suite for assessing VLA models. We profile three state-of-the-art VLM and VLAs - GPT-4o, OpenVLA, and JAT - across 20 diverse datasets from the Open-X-Embodiment collection, evaluating their performance on various manipulation tasks. Our analysis reveals several key insights: 1. current VLA models show significant variation in performance across different tasks and robot platforms, with GPT-4o demonstrating the most consistent performance through sophisticated prompt engineering, 2. all models struggle with complex manipulation tasks requiring multi-step planning, and 3. model performance is notably sensitive to action space characteristics and environmental factors. We release our evaluation framework and findings to facilitate systematic assessment of future VLA models and identify critical areas for improvement in the development of general purpose robotic systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。