构建交互式评测基准,评估多模态大模型在真实场景中的具身能力
EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents
- 设计328个任务覆盖125个三维场景,统一仿真框架支持多模态模型评测
- 模型在导航、社交互动等任务上与人类表现差距显著,平均得分仅达人类62%
- 适合研究具身智能、多模态推理的学者和开发者使用
多模态大语言模型(MLLMs)在具身智能领域展现出巨大潜力,但现有评测基准多基于静态图像或视频,难以评估其交互能力。而现有具身AI评测任务单一且缺乏多样性。为此,我们提出EmbodiedEval,一个面向具身任务的综合性交互式评测基准。该基准包含125个多样化3D场景中的328个独立任务,涵盖导航、物体交互、社交互动、属性问答与空间问答五大类,全面评估模型的具身能力。所有任务均经过严格筛选与标注,并在统一仿真与评估框架中执行。我们在该基准上评估了当前最先进的MLLMs,发现其在具身任务上的表现与人类水平存在显著差距(平均得分仅为人类的62%)。分析揭示了现有模型在具身认知方面的局限性,为后续研究提供方向。相关数据与仿真框架已开源:https://github.com/thunlp/EmbodiedEval。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have shown significant advancements, providing a promising future for embodied agents. Existing benchmarks for evaluating MLLMs primarily utilize static images or videos, limiting assessments to non-interactive scenarios. Meanwhile, existing embodied AI benchmarks are task-specific and not diverse enough, which do not adequately evaluate the embodied capabilities of MLLMs. To address this, we propose EmbodiedEval, a comprehensive and interactive evaluation benchmark for MLLMs with embodied tasks. EmbodiedEval features 328 distinct tasks within 125 varied 3D scenes, each of which is rigorously selected and annotated. It covers a broad spectrum of existing embodied AI tasks with significantly enhanced diversity, all within a unified simulation and evaluation framework tailored for MLLMs. The tasks are organized into five categories: navigation, object interaction, social interaction, attribute question answering, and spatial question answering to assess different capabilities of the agents. We evaluated the state-of-the-art MLLMs on EmbodiedEval and found that they have a significant shortfall compared to human level on embodied tasks. Our analysis demonstrates the limitations of existing MLLMs in embodied capabilities, providing insights for their future development. We open-source all evaluation data and simulation framework at https://github.com/thunlp/EmbodiedEval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。