对比四款视觉语言动作模型在真实机器人上的表现,揭示其优劣与适用场景。
Experiences from Benchmarking Vision-Language-Action Models for Robotic Manipulation
- 构建标准化评估框架,测试模型在精度、适应性和指令理解三方面表现
- π₀在分布外场景适应性最强,ACT在分布内最稳定,各有优劣
- 发现模型存在计算开销大、抓取失误等共性问题,指导实际部署选择
将基础模型应用于机器人领域,尤其是视觉-语言-动作(VLA)模型,有望实现通用操作。然而,系统性的真实世界评估和跨模型比较仍很匮乏。本文报告了我们在四个代表性VLA模型——ACT、OpenVLA-OFT、RDT-1B和π₀——上,于模拟环境及ALOHA Mobile平台上开展的四类操作任务的实证经验。我们建立了一个标准化评估框架,从三个关键维度衡量性能:(1) 准确性与效率(成功率和达成时间),(2) 在分布内、空间分布外、实例+空间分布外设置下的适应能力,(3) 语言指令遵循准确率。结果显示,π₀在分布外场景中表现出更强的适应性,而ACT在分布内具有更高的稳定性。进一步分析揭示了各模型在计算需求、数据扩展行为以及重复性失败模式(如近距抓取失败、过早释放、长序列状态漂移)上的差异。这些发现揭示了VLA模型架构在精度、泛化能力与部署成本之间的实际权衡,为真实机器人操作中选择和部署VLA模型提供了可操作的洞见。
原文摘要 · Abstract (English)
Foundation models applied in robotics, particularly \textbf{Vision--Language--Action (VLA)} models, hold great promise for achieving general-purpose manipulation. Yet, systematic real-world evaluations and cross-model comparisons remain scarce. This paper reports our \textbf{empirical experiences} from benchmarking four representative VLAs -- \textbf{ACT}, \textbf{OpenVLA--OFT}, \textbf{RDT-1B}, and \boldmath{$π_0$} -- across four manipulation tasks conducted in both simulation and on the \textbf{ALOHA Mobile} platform. We establish a \textbf{standardized evaluation framework} that measures performance along three key dimensions: (1) \textit{accuracy and efficiency} (success rate and time-to-success), (2) \textit{adaptability} across in-distribution, spatial out-of-distribution, and instance-plus-spatial out-of-distribution settings, and (3) \textit{language instruction-following accuracy}. Through this process, we observe that \boldmath{$π_0$} demonstrates superior adaptability in out-of-distribution scenarios, while \textbf{ACT} provides the highest stability in-distribution. Further analysis highlights differences in computational demands, data-scaling behavior, and recurring failure modes such as near-miss grasps, premature releases, and long-horizon state drift. These findings reveal practical trade-offs among VLA model architectures in balancing precision, generalization, and deployment cost, offering actionable insights for selecting and deploying VLAs in real-world robotic manipulation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。