arXiv:2410.05191cs.ROcs.AI2024-10被引 10

为视觉语言动作模型设计自动测试平台,提升机器人任务评估效率。

LADEV: A Language-Driven Testing and Evaluation Platform for Vision-Language-Action Models in Robotic Manipulation

  • 用自然语言自动生成仿真环境,免去手动配置。
  • 通过改写指令生成多样任务,测试模型泛化能力。
  • 支持批量测试,加速大规模模型评估。

基于大语言模型(LLMs)和视觉语言模型(VLMs)的发展,视觉语言动作(VLA)模型作为机器人操作任务的集成解决方案被提出。该模型以图像和自然语言指令为输入,直接生成机器人控制动作,显著提升决策能力与人机交互性。然而,由于其数据驱动特性及缺乏可解释性,保障VLA模型的有效性和鲁棒性面临挑战,亟需可靠测试与评估平台。为此,本文提出LADEV——一个专为评估VLA模型设计的综合性高效平台。首先,提出一种语言驱动方法,从自然语言输入自动生成仿真环境,减少人工调整,大幅提升测试效率;其次,引入改写机制,生成多样化自然语言任务指令以评估语言输入对模型的影响;最后,设计批处理方式实现大规模模型测试。在多个先进VLA模型上使用LADEV进行实验,结果表明该平台不仅显著提升测试效率,还建立了评估基准,推动更智能机器人系统的发展。

原文摘要 · Abstract (English)

Building on the advancements of Large Language Models (LLMs) and Vision Language Models (VLMs), recent research has introduced Vision-Language-Action (VLA) models as an integrated solution for robotic manipulation tasks. These models take camera images and natural language task instructions as input and directly generate control actions for robots to perform specified tasks, greatly improving both decision-making capabilities and interaction with human users. However, the data-driven nature of VLA models, combined with their lack of interpretability, makes the assurance of their effectiveness and robustness a challenging task. This highlights the need for a reliable testing and evaluation platform. For this purpose, in this work, we propose LADEV, a comprehensive and efficient platform specifically designed for evaluating VLA models. We first present a language-driven approach that automatically generates simulation environments from natural language inputs, mitigating the need for manual adjustments and significantly improving testing efficiency. Then, to further assess the influence of language input on the VLA models, we implement a paraphrase mechanism that produces diverse natural language task instructions for testing. Finally, to expedite the evaluation process, we introduce a batch-style method for conducting large-scale testing of VLA models. Using LADEV, we conducted experiments on several state-of-the-art VLA models, demonstrating its effectiveness as a tool for evaluating these models. Our results showed that LADEV not only enhances testing efficiency but also establishes a solid baseline for evaluating VLA models, paving the way for the development of more intelligent and advanced robotic systems.

机器人语言模型评估平台

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。