新基准暴露视觉语言动作模型依赖记忆而非理解的缺陷
LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization

- 构建四维扰动评估体系,模拟真实环境变化
- 原模型90%准确率在新设置下骤降至0.0%
- 适合关注模型泛化能力与评测公正性的研究者
LIBERO已成为评估视觉语言动作(VLA)模型的广泛基准,但其现有训练与评估设置存在缺陷,常导致性能估计虚高,阻碍公平比较。为此,我们提出LIBERO-PRO,一个扩展的LIBERO基准,系统性地在四个维度——操作物体、初始状态、任务指令和环境——引入合理扰动以评估模型表现。实验表明,尽管现有模型在标准LIBERO评估中达到90%以上准确率,但在我们的泛化设置下性能骤降至0.0%。这一显著差异暴露了模型对训练集中的动作序列和环境布局的机械记忆依赖,而非真正任务理解或环境感知。例如,当目标物体被替换为无关物品时,模型仍持续执行抓取动作;即使指令被破坏或变为乱码,输出也保持不变。这些发现揭示了当前评估方法的严重缺陷,我们呼吁学界摒弃误导性方法,转向稳健的模型泛化与理解能力评估。代码已开源:https://github.com/Zxy-MLlab/LIBERO-PRO。
原文摘要 · Abstract (English)
LIBERO has emerged as a widely adopted benchmark for evaluating Vision-Language-Action (VLA) models; however, its current training and evaluation settings are problematic, often leading to inflated performance estimates and preventing fair model comparison. To address these issues, we introduce LIBERO-PRO, an extended LIBERO benchmark that systematically evaluates model performance under reasonable perturbations across four dimensions: manipulated objects, initial states, task instructions, and environments. Experimental results reveal that, although existing models achieve over 90% accuracy under the standard LIBERO evaluation, their performance collapses to 0.0% under our generalized setting. Crucially, this discrepancy exposes the models' reliance on rote memorization of action sequences and environment layouts from the training set, rather than genuine task understanding or environmental perception. For instance, models persist in executing grasping actions when the target object is replaced with irrelevant items, and their outputs remain unchanged even when given corrupted instructions or even messy tokens. These findings expose the severe flaws in current evaluation practices, and we call on the community to abandon misleading methodologies in favor of robust assessments of model generalization and comprehension. Our code is available at: https://github.com/Zxy-MLlab/LIBERO-PRO.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。