arXiv:2508.13142cs.CVcs.CL2025-08被引 16

评测主流多模态模型的空间智能,发现最强模型仍远不如人类。

Holistic Evaluation of Multimodal LLMs on Spatial Intelligence

  • 构建统一空间任务体系EASI,整合多个基准与新数据集。
  • 测试耗时超十亿token,GPT-5在多数空间任务上仍落后人类。
  • 开源工具链与排行榜,推动空间智能研究协作与可复现。

近年来多模态模型取得显著进展,但在空间理解与推理能力上仍存在明显短板,而这一能力正是人工智能与物理世界连接的核心。随着传闻中的最强大模型GPT-5发布,有必要评估当前领先模型(GPT、Gemini、Grok、Seed、Qwen、Intern)在迈向空间智能(SI)过程中的实际水平。为此,我们提出EASI:面向多模态大模型空间智能的全面评估框架。EASI构建了一个涵盖现有基准与新增任务的综合性空间任务分类体系,支持对前沿模型进行系统性评测。本报告在八个关键基准上开展评估,总计算成本超过十亿个令牌。实证结果表明:(1)GPT-5在空间智能方面展现出前所未有的能力;(2)但在广泛的空间任务中仍显著落后于人类表现;(3)空间任务暴露出比非空间任务更严重的模型能力缺陷;(4)在最困难的任务上,专有模型并未体现出决定性优势。此外,我们在一系列对人类直观但对先进多模态模型却失败的场景中进行了定性分析。EASI是一项持续的社区努力:我们已开源EASI代码库,提供一站式、可复现的解决方案,包含标准化接口、集成协议和提示模板,极大降低多基准配置与运行的门槛;同时推出配套的EASI排行榜,持续更新模型在全谱系空间智能任务上的表现,加速向鲁棒空间智能的集体迈进。

原文摘要 · Abstract (English)

Multimodal models have achieved remarkable progress in recent years. Nevertheless, they continue to exhibit notable limitations in spatial understanding and reasoning, the very capability that anchors artificial general intelligence in the physical world. With the recent release of GPT-5, allegedly the most powerful AI model to date, it is timely to examine where the leading models (GPT, Gemini, Grok, Seed, Qwen, and Intern) stand on the path toward spatial intelligence (SI). We thus propose EASI for holistic Evaluation of multimodAl LLMs on Spatial Intelligence. EASI conceptualizes a comprehensive taxonomy of spatial tasks that unifies existing benchmarks and a growing collection of newly curated ones, enabling systematic evaluation of state-of-the-art models. In this report, we conduct the study across eight key benchmarks, at a cost exceeding ten billion total tokens. Our empirical study then reveals that (1) GPT-5 demonstrates unprecedented strength in SI, yet (2) still falls short of human performance significantly across a broad spectrum of SI-tasks. Moreover, we (3) show that SI-tasks expose greater model capability deficiency than non-SI tasks, to the extent that (4) proprietary models do not exhibit a decisive advantage when facing the most difficult ones. In addition, we conduct a qualitative evaluation across a diverse set of scenarios that are intuitive for humans, yet fail the most advanced multimodal models. EASI is an ongoing community effort: we have open-sourced the EASI codebase that provides a one-stop and reproducible solution with standardized interfaces, integrated protocols and prompts that significantly reduce the friction of configuring and running multiple benchmarks; we have also launched an accompanying EASI leaderboard to provide a continually updated snapshot of model performance across the full SI spectrum, accelerating collective progress toward robust SI.

多模态空间智能评测GPT-5

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。