构建首个空中-地面协同推理基准,评估视觉语言模型在真实场景中的表现
AeroGround: A Comprehensive Benchmark for Aerial-Ground Collaborative Reasoning

- 基于2.9万组多模态观测数据构建仿真数据集
- 16个预训练模型最高准确率54.4%,远低于人类的93.3%
- 聚焦跨视角对应、空间理解等真实任务,适合智能无人系统研究者
视觉语言模型(VLMs)已广泛用于无人机(UAV)的感知与推理任务。现有无人机基准主要关注空中视角场景,但当前模型在真实应用中常见的空中-地面协同任务(如救援、基础设施巡检)中的表现仍缺乏深入研究。为此,我们提出AeroGround,一个面向空中-地面协同推理的综合性评估基准。该基准基于包含约2.9万组多模态观测数据的仿真数据集,提供2,250个高质量问答实例,覆盖跨视角对应、空间理解与推理等任务。对16个预训练VLMs及两个领域自适应变体的实验显示,当前模型与人类性能存在显著差距:最佳模型平均准确率为54.4%,而人类达到93.3%。AeroGround系统揭示了现有模型在空中-地面协同推理中的优劣,为发展更强大的协同智能系统奠定基础。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have been widely employed in understanding and reasoning tasks for unmanned aerial vehicles (UAVs). Existing UAV benchmarks primarily focus on aerial-view scenarios. However, whether current VLMs can perform well on understanding and reasoning tasks in aerial-ground collaborative scenarios which are practical in real-world applications like rescue and infrastructure inspection remains underexplored. To address this gap, we introduce AeroGround, a comprehensive benchmark for evaluating VLMs in aerial-ground collaborative reasoning. AeroGround is built upon a simulated aerial-ground dataset containing approximately 29,000 multimodal observation groups from diverse open environments, and provides 2,250 high-quality question-answering instances covering cross-view correspondence, spatial understanding, and reasoning. Experiments on 16 pretrained VLMs, together with two domain-adapted variants, reveal a substantial gap between current models and human performance: the best model achieves an average accuracy of 54.4%, whereas humans reach 93.3%. By systematically revealing the strengths and limitations of existing models in aerial-ground collaborative reasoning, AeroGround provides a foundation for developing more capable aerial-ground collaborative embodied intelligence systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。