首个系统评估视觉语言动作模型抗物理扰动能力的框架,发现主流模型在真实场景中失败率超90%。
Eva-VLA: Evaluating Vision-Language-Action Models' Robustness Under Real-World Physical Variations
- 将物理扰动拆解为三维变换、光照变化和对抗区域三维度,构建可复现的评估体系
- 通过连续黑箱优化生成最坏情况,实测OpenVLA在LIBERO-Long任务上平均失败率达90%以上
- 生成的对抗样本可用于训练,显著提升模型鲁棒性,适合机器人研发与测评人员
视觉语言动作(VLA)模型在机器人操作中展现出巨大潜力,但其对真实世界物理变化的鲁棒性仍严重缺乏研究。为此,我们提出Eva-VLA,首个统一框架,将不可控物理变化建模为连续优化问题,系统评估VLA模型的鲁棒性。针对两大挑战:1)如何系统表征真实部署中多样的物理扰动并保持可复现性;2)如何高效发现最坏情况而不产生高昂的现实数据采集成本。针对第一个挑战,我们将现实变化分解为三个关键维度:影响空间推理的3D物体变换、挑战视觉感知的光照变化、破坏场景理解的对抗区域。针对第二个挑战,引入连续黑箱优化机制,将这些扰动映射到连续参数空间,实现最坏情况的系统探索。大量实验验证了该方法的有效性。值得注意的是,OpenVLA在LIBERO-Long任务上面对三种物理变化时平均失败率超过90%,暴露出严重的系统脆弱性。此外,利用生成的最坏情况样本进行对抗训练可量化提升模型鲁棒性,验证了该方法的有效性。评估揭示了实验室与真实环境间的差距,而Eva-VLA框架本身可作为有效数据增强手段,提升机器人操作系统的韧性。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have emerged as promising solutions for robotic manipulation, yet their robustness to real-world physical variations remains critically underexplored. To bridge this gap, we propose Eva-VLA, the first unified framework to systematically evaluate the robustness of VLA models by formulating uncontrollable physical variations as continuous optimization problems. Specifically, our framework addresses two fundamental challenges in VLA models' physical robustness evaluation: 1) how to systematically characterize diverse physical perturbations encountered in real-world deployment while maintaining reproducibility, and 2) how to efficiently discover worst-case scenarios without incurring prohibitive real-world data collection costs. To tackle the first challenge, we decouple real-world variations into three key dimensions: 3D object transformations that affect spatial reasoning, illumination changes that challenge visual perception, and adversarial regions that disrupt scene understanding. For the second challenge, we introduce a continuous black-box optimization mechanism that maps these perturbations into a continuous parameter space, enabling the systematic exploration of worst-case scenarios. Extensive experiments validate the effectiveness of our approach. Notably, OpenVLA exhibits an average failure rate of over 90% across three physical variations on the LIBERO-Long task, exposing critical systemic fragilities. Furthermore, applying the generated worst-case scenarios during adversarial training quantifiably increases model robustness, validating the effectiveness of this approach. Our evaluation exposes the gap between laboratory and real-world conditions, while the Eva-VLA framework can serve as an effective data augmentation method to enhance the resilience of robotic manipulation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。