arXiv:2603.28301cs.LG2026-03中稿 · EMNLP被引 5

提出新基准与评估方法,揭示视觉语言动作模型对指令改写敏感性。

LIBERO-Para: A Diagnostic Benchmark and Metrics for Paraphrase Robustness in VLA Models

  • 构建独立控制动作与物体表达的精细指令改写基准
  • 模型在指令改写下性能下降22-52个百分点,主要因物体词汇变化
  • 提出PRIDE度量法,区分改写难度,揭示模型依赖表面匹配而非语义理解

视觉语言动作(VLA)模型通过预训练视觉语言主干实现强机器人操作性能。然而,在下游任务中通常仅用少量数据微调,导致对特定指令表述过拟合,对指令改写鲁棒性研究不足。为此,我们提出LIBERO-Para,一个可控制的基准,独立调节动作表达与物体引用,实现语言泛化能力的细粒度分析。在七个VLA配置(0.6B-7.5B)中,我们观察到指令改写下性能一致下降22-52个百分点。该下降主要由物体层面的词汇变化驱动:即使简单同义替换也引发显著性能下滑,表明模型依赖表层匹配而非语义接地。此外,80-96%的失败源于规划级轨迹偏差,而非执行错误,说明改写干扰了任务识别。二元成功率将所有改写等同对待,掩盖了模型在不同难度下的表现差异或对较易情况的依赖。为此,我们提出PRIDE,一个基于语义与句法因素量化改写难度的指标。相关基准与代码已公开于:https://github.com/cau-hai-lab/LIBERO-Para

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models achieve strong performance in robotic manipulation by leveraging pre-trained vision-language backbones. However, in downstream robotic settings, they are typically fine-tuned with limited data, leading to overfitting to specific instruction formulations and leaving robustness to paraphrased instructions underexplored. To study this gap, we introduce LIBERO-Para, a controlled benchmark that independently varies action expressions and object references for fine-grained analysis of linguistic generalization. Across seven VLA configurations (0.6B-7.5B), we observe consistent performance degradation of 22-52 pp under paraphrasing. This degradation is primarily driven by object-level lexical variation: even simple synonym substitutions cause large drops, indicating reliance on surface-level matching rather than semantic grounding. Moreover, 80-96% of failures arise from planning-level trajectory divergence rather than execution errors, showing that paraphrasing disrupts task identification. Binary success rate treats all paraphrases equally, obscuring whether models perform consistently across difficulty levels or rely on easier cases. To address this, we propose PRIDE, a metric that quantifies paraphrase difficulty using semantic and syntactic factors. Our benchmark and corresponding code are available at: https://github.com/cau-hai-lab/LIBERO-Para

视觉语言动作指令鲁棒性评估基准机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。