arXiv:2509.15273cs.RO2025-09被引 10

构建统一演进的具身智能评估平台,解决能力体系与评测标准缺失问题。

Embodied Arena: A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AI

  • 建立三层七类二十五条具身能力分类体系,明确研究目标。
  • 集成22个基准、30+模型,支持跨任务多维度实时评测。
  • 用大模型自动生成数据,实现评估数据持续进化与扩展。

具身智能发展因三大挑战滞后:(1)缺乏对核心能力的系统理解,研究目标不清晰;(2)缺少统一标准化评测体系,跨基准比较不可行;(3)具身数据自动化获取方法薄弱,制约模型扩展。为此,我们提出Embodied Arena——一个全面、统一且可演进的具身智能评估平台。该平台建立涵盖感知、推理、任务执行三个层级,七项核心能力、25个细粒度维度的系统性能力分类体系,支持统一评估与明确研究方向。构建基于统一基础设施的标准化评测系统,灵活整合22个不同基准,覆盖2D/3D具身问答、导航、任务规划三大领域,支持30+来自20+机构的先进模型。开发基于大语言模型的自动化生成管道,实现可扩展、持续演进的具身评估数据。发布三个实时排行榜(具身问答、导航、任务规划),提供基准视角与能力视角双维度视图,全面呈现模型性能。基于排行榜结果总结出九项关键发现,为研究指明方向并定位关键问题,推动具身智能领域进展。

原文摘要 · Abstract (English)

Embodied AI development significantly lags behind large foundation models due to three critical challenges: (1) lack of systematic understanding of core capabilities needed for Embodied AI, making research lack clear objectives; (2) absence of unified and standardized evaluation systems, rendering cross-benchmark evaluation infeasible; and (3) underdeveloped automated and scalable acquisition methods for embodied data, creating critical bottlenecks for model scaling. To address these obstacles, we present Embodied Arena, a comprehensive, unified, and evolving evaluation platform for Embodied AI. Our platform establishes a systematic embodied capability taxonomy spanning three levels (perception, reasoning, task execution), seven core capabilities, and 25 fine-grained dimensions, enabling unified evaluation with systematic research objectives. We introduce a standardized evaluation system built upon unified infrastructure supporting flexible integration of 22 diverse benchmarks across three domains (2D/3D Embodied Q&A, Navigation, Task Planning) and 30+ advanced models from 20+ worldwide institutes. Additionally, we develop a novel LLM-driven automated generation pipeline ensuring scalable embodied evaluation data with continuous evolution for diversity and comprehensiveness. Embodied Arena publishes three real-time leaderboards (Embodied Q&A, Navigation, Task Planning) with dual perspectives (benchmark view and capability view), providing comprehensive overviews of advanced model capabilities. Especially, we present nine findings summarized from the evaluation results on the leaderboards of Embodied Arena. This helps to establish clear research veins and pinpoint critical research problems, thereby driving forward progress in the field of Embodied AI.

具身智能评估平台多模态自动化评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。