NEBULA新评估体系揭示视觉语言动作模型真实短板
NEBULA: Do We Evaluate Vision-Language-Action Agents Correctly?
- 设计双轴评测机制,兼顾精细技能诊断与抗干扰能力测试
- 实测顶尖模型在空间推理和动态适应上表现不佳
- 提供统一接口与数据集,助力可复现研究与模型对比
视觉语言动作(VLA)代理的评估受限于粗粒度的最终任务成功率指标,难以精确诊断技能水平或衡量对现实扰动的鲁棒性。这一问题因数据碎片化而加剧,阻碍了可复现研究和通用模型的发展。为此,我们提出NEBULA,一个面向单臂操作的统一生态系统,支持诊断性与可复现的评估。NEBULA包含新颖的双轴评估协议:细粒度能力测试用于精准技能诊断,系统性压力测试用于衡量鲁棒性。同时提供标准化API和大规模聚合数据集,减少数据碎片化,支持跨数据集训练与公平比较。利用NEBULA,我们发现顶级VLA在空间推理和动态适应等关键能力上表现不佳,这些缺陷被传统终点成功度量所掩盖。通过同时测量能力与可靠性,NEBULA为构建鲁棒、通用的具身智能体提供了实践基础。
原文摘要 · Abstract (English)
The evaluation of Vision-Language-Action (VLA) agents is hindered by the coarse, end-task success metric that fails to provide precise skill diagnosis or measure robustness to real-world perturbations. This challenge is exacerbated by a fragmented data landscape that impedes reproducible research and the development of generalist models. To address these limitations, we introduce NEBULA, a unified ecosystem for single-arm manipulation that enables diagnostic and reproducible evaluation. NEBULA features a novel dual-axis evaluation protocol that combines fine-grained capability tests for precise skill diagnosis with systematic stress tests that measure robustness. A standardized API and a large-scale, aggregated dataset are provided to reduce fragmentation and support cross-dataset training and fair comparison. Using NEBULA, we demonstrate that top-performing VLAs struggle with key capabilities such as spatial reasoning and dynamic adaptation, which are consistently obscured by conventional end-task success metrics. By measuring both what an agent can do and when it does so reliably, NEBULA provides a practical foundation for robust, general-purpose embodied agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。