评测无人机与地面机器人协作中的多视角空间理解能力。
AirGroundBench: Probing Spatial Intelligence in Multimodal Large Models under Heterogeneous Multi-View Embodied Collaboration

- 构建11个仿真环境,含1021对同步空地观测数据。
- 13个模型在双视角下仍存在跨视图对齐和推理短板。
- 适合研究多智能体协作与空间认知的学者参考。
近年来,多模态大语言模型(MLLMs)在具身智能方面展现出巨大潜力,但其在异构视角下保持几何一致性空间理解的能力尚未得到充分评估。现有基准主要聚焦单智能体、单视角感知,缺乏对空地协同场景中多尺度观测互补性带来的尺度不匹配、非对称遮挡和参考系不一致等问题的系统评估。本文提出AirGroundBench,一个用于诊断异构无人机-地面机器人协作中多视角空间智能的基准。该基准基于11个高保真仿真环境,包含1,021对同步空地观测数据,生成约62,000个双视角四选一视觉问答实例及115个闭环视觉-语言导航任务。涵盖10类任务,分为四个逐步递增的能力维度:空间感知、跨视角对齐、空间变换与推理、具身决策。为支持几何基础评估,提供结构化空间标注,包括跨视角物体身份及度量2D/3D边界框。在13个代表性MLLM上进行无人机独用、地面机器人独用及双视角输入设置下的评估显示,模型在空间感知方面表现尚可,但在跨视角对齐与变换密集型推理上存在明显缺陷,且影响序列决策。尽管双视角输入带来可测量提升,但与人类表现仍有显著差距,凸显几何一致性是当前具身MLLM的关键瓶颈。
原文摘要 · Abstract (English)
In recent years, multimodal large language models (MLLMs) have shown strong potential for embodied intelligence, yet their ability to maintain geometrically consistent spatial understanding across heterogeneous views remains under-evaluated. Existing benchmarks largely focus on single-agent, single-view perception, leaving a gap in the systematic assessment of collaborative air-ground settings, where multi-scale observations are complementary but introduce scale mismatch, asymmetric occlusion, and reference-frame inconsistencies. We present AirGroundBench, a diagnostic benchmark for evaluating multi-view spatial intelligence in heterogeneous UAV-UGV collaboration. AirGroundBench is built from 11 high-fidelity simulated environments with 1,021 synchronized air-ground observation pairs, yielding approximately 62,000 dual-view, four-option single-choice visual question answering instances and 115 closed-loop vision-language navigation episodes. It covers 10 task types organized into four progressively demanding capability dimensions: spatial perception, cross-view alignment, spatial transformation and reasoning, and embodied decision-making. To support geometry-grounded evaluation and analysis, we provide structured spatial annotations, including cross-view object identities and metric 2D and 3D bounding boxes. Evaluations of 13 representative MLLMs under UAV-only, UGV-only, and dual-view input settings reveal consistent bottlenecks: models perform relatively well on spatial perception but struggle with cross-view alignment and transformation-intensive reasoning, and these deficits propagate to sequential decision-making in vision-language navigation. Although dual-view inputs provide measurable gains over single-view variants, a persistent gap from human performance remains, highlighting geometric consistency as a key limitation of current embodied MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。