arXiv:2602.06556cs.CVcs.AI2026-02被引 10

新基准LIBERO-X测试视觉语言动作模型在复杂环境下的鲁棒性。

LIBERO-X: Robustness Litmus for Vision-Language-Action Models

  • 分层级评估体系,逐步增加任务难度
  • 实测多模型在扰动下性能显著下降
  • 适合研究模型泛化与指令理解的学者

可靠的基准测试对推动视觉-语言-动作(VLA)模型发展至关重要,可揭示其泛化能力、鲁棒性及感知与语言驱动操作任务的对齐程度。然而,现有基准因评估协议不足,难以真实反映分布偏移问题,常导致评估结果有限或误导。本文从评估与数据双视角重构VLA基准,提出LIBERO-X:1)采用分层评估协议,设置渐进式难度等级,聚焦空间泛化、物体识别和任务指令理解三大核心能力,实现对环境与任务复杂度增加时性能退化的细粒度分析;2)通过人工遥操作收集高多样性训练数据集,每个场景支持多个精细操作目标,缩小训练与评估分布差距。代表性VLA模型实验显示,在累积扰动下性能显著下降,暴露出场景理解与指令定位的持续局限。结合分层评估与多样化数据,LIBERO-X为更可靠地评估与推进VLA发展提供了坚实基础。

原文摘要 · Abstract (English)

Reliable benchmarking is critical for advancing Vision-Language-Action (VLA) models, as it reveals their generalization, robustness, and alignment of perception with language-driven manipulation tasks. However, existing benchmarks often provide limited or misleading assessments due to insufficient evaluation protocols that inadequately capture real-world distribution shifts. This work systematically rethinks VLA benchmarking from both evaluation and data perspectives, introducing LIBERO-X, a benchmark featuring: 1) A hierarchical evaluation protocol with progressive difficulty levels targeting three core capabilities: spatial generalization, object recognition, and task instruction understanding. This design enables fine-grained analysis of performance degradation under increasing environmental and task complexity; 2) A high-diversity training dataset collected via human teleoperation, where each scene supports multiple fine-grained manipulation objectives to bridge the train-evaluation distribution gap. Experiments with representative VLA models reveal significant performance drops under cumulative perturbations, exposing persistent limitations in scene comprehension and instruction grounding. By integrating hierarchical evaluation with diverse training data, LIBERO-X offers a more reliable foundation for assessing and advancing VLA development.

视觉语言动作模型评估鲁棒性测试机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。