首个面向视觉驱动具身智能体的多模态大模型评测基准
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

- 构建涵盖1128个任务的跨环境评测框架,覆盖高阶语义与低阶操作
- 顶尖模型GPT-4o在低阶操作任务上平均仅达28.9%准确率
- 适合研究具身智能、多模态推理与机器人交互的学者使用
利用多模态大语言模型(MLLMs)构建具身智能体为解决现实任务提供了新路径。尽管以语言为中心的具身智能体已受广泛关注,但基于MLLM的具身智能体因缺乏全面评估框架而研究不足。为此,我们提出EmbodiedBench,一个全面的评测基准,用于评估视觉驱动的具身智能体。该基准包含:(1) 四个环境中1,128个测试任务,涵盖从高阶语义任务(如家庭场景)到低阶原子动作任务(如导航与操作);(2) 六个精心设计的子集,评估常识推理、复杂指令理解、空间感知、视觉感知和长期规划等核心能力。我们对24个主流专有及开源MLLM进行了广泛实验。结果表明:MLLM在高阶任务表现优异,但在低阶操作上仍显著不足,最优模型GPT-4o平均得分仅为28.9%。EmbodiedBench提供了一个多维度标准化评估平台,不仅揭示了现有挑战,也为推进MLLM驱动的具身智能体发展提供了宝贵洞见。代码与数据集见https://embodiedbench.github.io。
原文摘要 · Abstract (English)
Leveraging Multi-modal Large Language Models (MLLMs) to create embodied agents offers a promising avenue for tackling real-world tasks. While language-centric embodied agents have garnered substantial attention, MLLM-based embodied agents remain underexplored due to the lack of comprehensive evaluation frameworks. To bridge this gap, we introduce EmbodiedBench, an extensive benchmark designed to evaluate vision-driven embodied agents. EmbodiedBench features: (1) a diverse set of 1,128 testing tasks across four environments, ranging from high-level semantic tasks (e.g., household) to low-level tasks involving atomic actions (e.g., navigation and manipulation); and (2) six meticulously curated subsets evaluating essential agent capabilities like commonsense reasoning, complex instruction understanding, spatial awareness, visual perception, and long-term planning. Through extensive experiments, we evaluated 24 leading proprietary and open-source MLLMs within EmbodiedBench. Our findings reveal that: MLLMs excel at high-level tasks but struggle with low-level manipulation, with the best model, GPT-4o, scoring only 28.9\% on average. EmbodiedBench provides a multifaceted standardized evaluation platform that not only highlights existing challenges but also offers valuable insights to advance MLLM-based embodied agents. Our code and dataset are available at https://embodiedbench.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。