arXiv:2605.27759cs.RO2026-05被引 1

构建大规模仿真基准,评估视觉语言动作模型在多种条件下的泛化能力。

Colosseum V2: Benchmarking Generalization for Vision Language Action Models

论文配图:Colosseum V2: Benchmarking Generalization for Vision Language Action Models
图 1 · 摘自论文原文
  • 基于ManiSkill搭建28个任务的统一仿真平台,支持跨域测试。
  • 实测ACT和Pi0.5模型在分布外场景下性能显著下降。
  • 仿真结果与真实机器人表现高度相关,适合研究通用机器人策略。

视觉-语言-动作(VLA)模型在机器人操作中展现出良好的泛化能力,这得益于大规模视觉与语言预训练的发展。然而,这种进步可能具有误导性:尽管VLA具备零样本感知和语言理解能力,其整体任务表现仍会在分布偏移下显著退化,暴露出高层理解向鲁棒行为转化的不足。为此,我们提出Colosseum V2,一个大规模仿真基准,用于系统评估机器人学习中VLA的泛化能力。该基准包含28个任务,覆盖13类任务和两种机器人形态,涵盖多样化的操作基本动作与长时序行为。基于ManiSkill仿真器,支持快速、GPU并行评估,并可实现域内与域外测试。我们评估了当前先进方法,包括Action Chunking Transformers(ACT)和Pi0.5,揭示其基础性能与泛化能力的局限性。实验表明,仿真结果与真实世界指标高度相关,验证了基准的生态有效性。通过统一任务、度量与评估协议,Colosseum V2实现了可复现、公平的对比,降低评估开销,加速通用机器人策略的研究进程。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models demonstrate promising generalization in robotic manipulation, driven by advances in large-scale vision and language pre-training. This progress can be misleading. Despite the zero-shot perception and language capabilities of VLAs, their overall task performance often degrades under distribution shifts, revealing gaps in how these systems translate high-level understanding into robust behavior. To systematically study this gap, we introduce Colosseum V2, a large-scale simulation benchmark for evaluating VLA generalization in robot learning across diverse conditions. The benchmark comprises 28 tasks spanning 13 task categories and two robot morphologies, covering a wide range of manipulation primitives and long-horizon behaviors. Built on the ManiSkill simulator, Colosseum V2 enables fast, GPU-parallelized evaluation and supports both in-domain and out-of-domain testing at scale. We evaluate state-of-the-art methods, including Action Chunking Transformers (ACT) and Pi0.5, and reveal limitations in both base performance and generalization. We demonstrate strong correlations between simulation and real-world metrics that support the ecological validity of the benchmark. By standardizing tasks, metrics, and evaluation protocols within a unified benchmark, Colosseum V2 enables reproducible and fair comparisons, reduced evaluation overhead, and accelerated progress toward general-purpose robot policies.

机器人学习泛化能力仿真基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。