低成本可复现的机器人视觉语言动作模型真实世界评测基准
VLA-REPLICA: A Low-Cost, Reproducible Benchmark for Real-World Evaluation of Vision-Language-Action Models

- 用市售零件搭建可快速复制的评测系统
- 支持多种操作任务与分布内外测试,结果跨实验室一致
- 适合研究者评估真实场景下VLA模型性能
视觉-语言-动作(VLA)模型在通用机器人操作中展现出巨大潜力,但其真实世界评估受限于缺乏低成本、可复现且一致的基准。仿真基准无法捕捉真实复杂性,现有真实世界基准往往需要昂贵硬件、集中式评估或任务多样性不足。我们提出VLA-REPLICA,一个基于市售组件构建的低成本、易复现的真实世界评测基准。该系统可快速组装并在各实验室复刻,提供全球一致的策略评估环境。包含多样化的操作任务和小规模示范数据集用于目标域适应,支持分布内与分布外评估协议。通过模仿学习与先进VLA模型的实验,揭示了模型的优势与局限,且独立构建系统间结果一致,验证了基准的可复现性。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have shown strong promise for general-purpose robotic manipulation, but their real-world evaluation remains limited by a lack of accessible, reproducible, and consistent benchmarks. Simulation benchmarks fail to capture real-world complexity, while existing real-world benchmarks often require expensive hardware, centralized evaluation, or are limited in task diversity. We introduce VLA-REPLICA, a low-cost, easily reproducible real-world benchmark for evaluating VLA models. Built from off-the-shelf components, our system can be quickly assembled and replicated across laboratories, providing a consistent environment for policy evaluation anywhere in the world. VLA-REPLICA includes a diverse suite of manipulation tasks and a small-scale demonstration dataset for target-domain adaptation, with real-world evaluation protocols for both in-distribution and out-of-distribution settings. Experiments with imitation learning and state-of-the-art VLA models reveal model strengths and limitations, while consistent results across independently constructed setups demonstrate the reproducibility of our benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。