开源评测框架,系统测试视觉-语言-动作模型的泛化与安全能力。
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
- 按任务结构、语言指令、视觉观察三轴设计170个渐进难度任务
- 发现顶尖模型存在记忆替代泛化、忽视安全约束等关键缺陷
- 适合研究机器人通用策略与模型鲁棒性的学者使用
尽管视觉-语言-动作模型(VLAs)正快速向通用机器人策略发展,但量化评估其能力边界与失效模式仍具挑战。为此,我们提出VLA-Arena,一个全面的基准测试框架。其采用新型结构化任务设计,从任务结构、语言指令和视觉观察三个正交维度量化难度,支持精细分级的任务构建,精准刻画模型能力前沿。任务结构包含11个任务套件,分安全、干扰、外推和长时程四维,共170个任务,每套覆盖L0-L2三级难度,仅允许在L0级微调以严格检验泛化能力。同时,语言(W0-W4)与视觉(V0-V4)扰动可独立应用于任意任务,作为诊断探针区分稳健理解与表面模式匹配。对现有先进VLAs的广泛评估揭示了关键局限:过度依赖记忆而非泛化、浅层视觉感知、忽视安全约束。不同难度层级(L0-L2)间的模型排名反转证实各层级提供非冗余信息。为推动研究并保障可复现性,我们公开完整框架,含端到端工具链及VLA-Arena-S/M/L数据集用于微调。基准、数据集、模型与排行榜已开放于https://vla-arena.github.io。
原文摘要 · Abstract (English)
While Vision-Language-Action models (VLAs) are rapidly advancing toward generalist robot policies, quantitatively characterizing their capability boundaries and failure modes remains challenging. To address this, we introduce VLA-Arena, a comprehensive benchmark. It features a novel structured task design framework to quantify difficulty across three orthogonal axes: (1) Task Structure, (2) Language Command, and (3) Visual Observation. This allows us to systematically design tasks with fine-grained difficulty levels, enabling a precise measurement of model capability frontiers. For task structure, VLA-Arena comprises 11 task suites organized into four dimensions: Safety, Distractor, Extrapolation, and Long Horizon, totaling 170 tasks. Each suite spans three difficulty levels (L0-L2), with fine-tuning restricted to L0 to rigorously assess generalization. Orthogonal to this, language (W0-W4) and visual (V0-V4) perturbations can be applied to any task as diagnostic probes to distinguish robust grounding from superficial pattern matching. Our extensive evaluation of state-of-the-art VLAs reveals critical limitations: memorization over generalization, superficial visual perception, and a neglect of safety constraints. Additionally, model rank reversals across L0-L2 validate that each level provides non-redundant insights. To foster research addressing these model limitations and ensure reproducibility, we provide the complete VLA-Arena framework, including an end-to-end toolchain from task definition to automated evaluation and the VLA-Arena-S/M/L datasets for fine-tuning. Our benchmark, datasets, models, and leaderboard are publicly available at https://vla-arena.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。