用普通航拍图评估AI对三维地理信息的理解能力
Geo3DVQA: Evaluating Vision-Language Models for 3D Geospatial Reasoning from Aerial Imagery
- 基于RGB影像构建11万条问答数据,测试模型的3D空间推理能力
- 发现现有模型在高度感知和多特征融合上存在明显短板
- 适合关注地理智能、遥感分析与视觉语言模型的研究者
三维地理空间分析在城市规划、气候适应和环境评估中至关重要。然而,当前方法依赖昂贵的专用传感器(如LiDAR和多光谱传感器),限制了全球应用。现有基于传感器或规则的方法难以处理需整合多维3D线索、多样化查询及可解释推理的任务。本文提出Geo3DVQA,一个全面的基准,用于评估视觉语言模型(VLMs)仅从RGB影像进行高程感知的3D地理空间推理能力。不同于传统传感器框架,Geo3DVQA强调融合高程、天空视野因子与地表覆盖模式的真实场景。该基准包含16类任务中的11万条精心标注的问答对,涵盖单特征推断、多特征推理和应用级分析。对十种先进VLMs的系统评估揭示了其在从RGB图像到3D空间推理中的根本局限。结果表明,领域特定的指令微调能持续提升各类任务的表现,包括高程感知与开放性、面向应用的推理。Geo3DVQA为基于RGB的3D地理空间推理提供了统一且可解释的评估框架,指出了可扩展3D空间分析的关键挑战与机遇。代码与数据已公开于https://github.com/mm1129/Geo3DVQA。
原文摘要 · Abstract (English)
Three-dimensional geospatial analysis is critical for applications in urban planning, climate adaptation, and environmental assessment. However, current methodologies depend on costly, specialized sensors, such as LiDAR and multispectral sensors, which restrict global accessibility. Additionally, existing sensor-based and rule-driven methods struggle with tasks requiring the integration of multiple 3D cues, handling diverse queries, and providing interpretable reasoning. We present Geo3DVQA, a comprehensive benchmark that evaluates vision-language models (VLMs) in height-aware 3D geospatial reasoning from RGB imagery alone. Unlike conventional sensor-based frameworks, Geo3DVQA emphasizes realistic scenarios integrating elevation, sky view factors, and land cover patterns. The benchmark comprises 110k curated question-answer pairs across 16 task categories, including single-feature inference, multi-feature reasoning, and application-level analysis. Through a systematic evaluation of ten state-of-the-art VLMs, we reveal fundamental limitations in RGB-to-3D spatial reasoning. Our results further show that domain-specific instruction tuning consistently enhances model performance across all task categories, including height-aware and open-ended, application-oriented reasoning. Geo3DVQA provides a unified, interpretable framework for evaluating RGB-based 3D geospatial reasoning and identifies key challenges and opportunities for scalable 3D spatial analysis. The code and data are available at https://github.com/mm1129/Geo3DVQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。