构建首个面向长程推理的通用语言控制机器人操作基准测试
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

- 设计100类任务,2000+物体,支持强随机化与多步推理
- 评估视觉-语言-动作模型在常识、物理规律等能力上的表现
- 适合研究通用机器人智能与大模型融合的学者使用
通用具身智能体需理解自然语言指令并精准执行通用任务。近年来,基于基础模型的视觉-语言-动作模型(VLAs)在语言控制操作(LCM)任务上展现出巨大潜力。然而,现有基准难以满足VLAs及相应算法的需求。为此,我们提出VLABench,一个开源的通用LCM任务评估基准。该基准包含100个精心设计的任务类别,每类任务具有强随机性,覆盖2000多个物体。VLABench在四个关键方面优于以往基准:1)需世界知识与常识迁移的任务;2)含隐含人类意图的自然语言指令而非模板;3)要求多步推理的长程任务;4)同时评估动作策略与语言模型能力。基准涵盖对网格/纹理理解、空间关系、语义指令、物理规律、知识迁移与推理等多项能力的评估。为支持下游微调,我们通过融合启发式技能与先验信息的自动化框架生成高质量训练数据。实验表明,当前最先进的预训练VLAs及基于视觉语言模型的工作流在本基准任务中仍面临挑战。
原文摘要 · Abstract (English)
General-purposed embodied agents are designed to understand the users' natural instructions or intentions and act precisely to complete universal tasks. Recently, methods based on foundation models especially Vision-Language-Action models (VLAs) have shown a substantial potential to solve language-conditioned manipulation (LCM) tasks well. However, existing benchmarks do not adequately meet the needs of VLAs and relative algorithms. To better define such general-purpose tasks in the context of LLMs and advance the research in VLAs, we present VLABench, an open-source benchmark for evaluating universal LCM task learning. VLABench provides 100 carefully designed categories of tasks, with strong randomization in each category of task and a total of 2000+ objects. VLABench stands out from previous benchmarks in four key aspects: 1) tasks requiring world knowledge and common sense transfer, 2) natural language instructions with implicit human intentions rather than templates, 3) long-horizon tasks demanding multi-step reasoning, and 4) evaluation of both action policies and language model capabilities. The benchmark assesses multiple competencies including understanding of mesh\&texture, spatial relationship, semantic instruction, physical laws, knowledge transfer and reasoning, etc. To support the downstream finetuning, we provide high-quality training data collected via an automated framework incorporating heuristic skills and prior information. The experimental results indicate that both the current state-of-the-art pretrained VLAs and the workflow based on VLMs face challenges in our tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。