arXiv:2510.13626cs.ROcs.CL2025-10被引 223

测试视觉语言动作模型在7种扰动下的鲁棒性,发现性能暴跌至30%以下。

LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

论文配图:LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
图 1 · 摘自论文原文
  • 通过7类可控扰动系统评估模型鲁棒性
  • 摄像头视角变化使成功率从95%降至30%以下
  • 模型对语言指令基本无视,适合关注真实场景可靠性的人

视觉-语言-动作(VLA)模型在机器人操作基准测试中表现优异,但这些结果可能掩盖了其内在的脆弱性。我们通过在物体布局、相机视角、机器人初始状态、语言指令、光照条件、背景纹理和传感器噪声七个维度引入受控扰动,对多种前沿模型进行了系统性脆弱性分析。结果显示,尽管表面表现良好,但模型普遍存在严重脆弱性:对相机视角和机器人初始状态等扰动极为敏感,性能在适度扰动下从95%骤降至30%以下。令人意外的是,模型对语言指令变化几乎不敏感,进一步实验表明它们往往完全忽略语言指令。这些发现挑战了高基准分数等于真正能力的假设,强调需要更贴近现实变化的评估方式。

原文摘要 · Abstract (English)

Visual-Language-Action (VLA) models report impressive success rates on robotic manipulation benchmarks, yet these results may mask fundamental weaknesses in robustness. We perform a systematic vulnerability analysis by introducing controlled perturbations across seven dimensions: objects layout, camera viewpoints, robot initial states, language instructions, light conditions, background textures and sensor noise. We comprehensively analyzed multiple state-of-the-art models and revealed consistent brittleness beneath apparent competence. Our analysis exposes critical weaknesses: models exhibit extreme sensitivity to perturbation factors, including camera viewpoints and robot initial states, with performance dropping from 95% to below 30% under modest perturbations. Surprisingly, models are largely insensitive to language variations, with further experiments revealing that models tend to ignore language instructions completely. Our findings challenge the assumption that high benchmark scores equate to true competency and highlight the need for evaluation practices that assess reliability under realistic variation.

机器人鲁棒性VLA模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。