arXiv:2409.12894cs.SEcs.RO2024-09被引 45

提出测试框架VLATest,评估视觉语言动作模型在复杂场景下的鲁棒性。

VLATest: Testing and Evaluating Vision-Language-Action Models for Robotic Manipulation

  • 用模糊测试生成多样机器人操作场景,自动探测模型弱点。
  • 7个主流VLA模型在复杂环境下成功率普遍低于40%。
  • 适合关注机器人泛化能力与安全评估的研究者。

生成式AI和多模态基础模型的快速发展为机器人操作带来了新机遇。视觉语言动作(VLA)模型通过利用大规模视觉-语言数据和机器人示范,展现出强大的视觉运动控制潜力。然而,现有VLA模型通常仅在有限的手工设计场景中进行评估,其在多样化环境中的泛化性能与鲁棒性仍不清楚。为此,我们提出了VLATest——一个用于生成机器人操作场景的模糊测试框架,以系统性地测试VLA模型。基于该框架,我们对七个代表性VLA模型进行了实证研究。结果表明,当前VLA模型在实际部署中缺乏必要鲁棒性。我们进一步分析了干扰物数量、光照条件、相机姿态、未见物体以及任务指令变异等因素对模型性能的影响。研究揭示了现有VLA模型的局限性,强调需开展更深入研究以实现可靠可信的VLA应用。

原文摘要 · Abstract (English)

The rapid advancement of generative AI and multi-modal foundation models has shown significant potential in advancing robotic manipulation. Vision-language-action (VLA) models, in particular, have emerged as a promising approach for visuomotor control by leveraging large-scale vision-language data and robot demonstrations. However, current VLA models are typically evaluated using a limited set of hand-crafted scenes, leaving their general performance and robustness in diverse scenarios largely unexplored. To address this gap, we present VLATest, a fuzzing framework designed to generate robotic manipulation scenes for testing VLA models. Based on VLATest, we conducted an empirical study to assess the performance of seven representative VLA models. Our study results revealed that current VLA models lack the robustness necessary for practical deployment. Additionally, we investigated the impact of various factors, including the number of confounding objects, lighting conditions, camera poses, unseen objects, and task instruction mutations, on the VLA model's performance. Our findings highlight the limitations of existing VLA models, emphasizing the need for further research to develop reliable and trustworthy VLA applications.

机器人VLA模型测试框架鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。