arXiv:2602.15918cs.CVcs.AI2026-02被引 2

构建首个面向地球影像的空间推理评测基准,覆盖距离方向与拓扑关系定量分析。

EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery

  • 设计多模态大模型在地球影像上的空间推理评测框架,支持几何坐标输入
  • 包含超32.5万组问答对,涵盖距离、方向、拓扑关系等复杂推理任务
  • 适合研究地理智能、具身智能和多模态模型空间理解能力的学者使用

多模态大语言模型(MLLMs)的空间推理能力在计算机视觉领域日益受到关注,因其对需要精确交互物理世界的具身智能系统至关重要。然而,地球影像上的空间推理仍滞后,因其需将物体定位在地理参考图像中,并结合视觉线索与矢量几何坐标(如2D边界框、折线、多边形)进行距离、方向和拓扑关系的定量推理。现有地球影像基准主要聚焦于2D空间定位、图像描述和粗粒度空间关系(如简单方向或邻近提示),缺乏对定量方向与距离推理、系统性拓扑关系以及超出边界框的复杂对象几何的支持。为此,我们提出EarthSpatialBench,一个全面评估MLLMs在地球影像上空间推理能力的基准。该基准包含超过32.5万组问题-答案对,涵盖:(1) 关于空间距离与方向的定性与定量推理;(2) 系统性拓扑关系;(3) 单对象查询、对象对查询及复合聚合组查询;(4) 通过文本描述、视觉叠加和显式几何坐标(包括2D边界框、折线、多边形)表达的对象引用。我们在开源与专有模型上进行了广泛实验,揭示了当前MLLMs在空间推理方面的局限性。

原文摘要 · Abstract (English)

Benchmarking spatial reasoning in multimodal large language models (MLLMs) has attracted growing interest in computer vision due to its importance for embodied AI and other agentic systems that require precise interaction with the physical world. However, spatial reasoning on Earth imagery has lagged behind, as it uniquely involves grounding objects in georeferenced images and quantitatively reasoning about distances, directions, and topological relations using both visual cues and vector geometry coordinates (e.g., 2D bounding boxes, polylines, and polygons). Existing benchmarks for Earth imagery primarily focus on 2D spatial grounding, image captioning, and coarse spatial relations (e.g., simple directional or proximity cues). They lack support for quantitative direction and distance reasoning, systematic topological relations, and complex object geometries beyond bounding boxes. To fill this gap, we propose \textbf{EarthSpatialBench}, a comprehensive benchmark for evaluating spatial reasoning in MLLMs on Earth imagery. The benchmark contains over 325K question-answer pairs spanning: (1) qualitative and quantitative reasoning about spatial distance and direction; (2) systematic topological relations; (3) single-object queries, object-pair queries, and compositional aggregate group queries; and (4) object references expressed via textual descriptions, visual overlays, and explicit geometry coordinates, including 2D bounding boxes, polylines, and polygons. We conducted extensive experiments on both open-source and proprietary models to identify limitations in the spatial reasoning of MLLMs.

空间推理多模态模型地球影像地理智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。