arXiv:2605.25784cs.CVcs.MM2026-05

测试大模型能否用高程数据区分相似外观的自然场景。

VertiCue-Bench: Diagnosing Whether MLLMs Use Height Cues to Resolve 2D Ambiguity in Remote Sensing Natural Scenes

论文配图:VertiCue-Bench: Diagnosing Whether MLLMs Use Height Cues to Resolve 2D Ambiguity in Remote Sensing Natural Scenes
图 1 · 摘自论文原文
  • 构建首个基于高程图的遥感诊断基准,分离感知与推理环节。
  • 14个模型在联合约束任务中表现不如仅用可见光的基线。
  • 揭示大模型虽能读高程图但难转化为准确语义理解的短板。

多模态大语言模型(MLLMs)在地理空间推理方面取得显著进展,但现有遥感评测基准仍以二维为主,主要依赖光学外观。在自然环境中,由于光谱混淆严重,生态上不同的区域可能具有相似纹理,但在垂直结构上存在本质差异。此时,冠层高程模型(CHM)等三维结构数据成为语义消歧的关键几何证据。然而,当前模型是否真正能利用垂直线索解决外观模糊仍不明确。为此,我们提出VertiCue-Bench,首个基于CHM的地理空间推理诊断基准。该基准包含1,534个精心设计的实例,覆盖17项任务,明确分离低层级高度感知与语义推理。对14个先进通用及遥感专用MLLM的评估结合反事实模态测试,揭示出显著的感知-推理脱节现象:模型虽具备初步读取原始CHM高度线索的能力,却难以将其转化为可靠的语义推理,在需要联合约束的任务中表现甚至低于仅使用RGB的基线。总体而言,VertiCue-Bench暴露了自然场景理解中几何到语义的关键鸿沟,为提升地理空间MLLM提供了可操作的洞见。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have recently shown promising progress in geospatial reasoning. However, existing remote sensing benchmarks remain largely 2D-centric, evaluating models primarily on optical appearance. In natural environments, this paradigm breaks down due to severe spectral confusion, where ecologically distinct regions share similar textures but differ fundamentally in vertical structure. In such cases, explicit 3D structural data, such as Canopy Height Models (CHMs), become essential geometric evidence for semantic disambiguation. Yet, it remains unclear whether current MLLMs can genuinely leverage vertical cues to resolve appearance-level ambiguity. To address this gap, we introduce VertiCue-Bench, the first diagnostic benchmark for CHM-grounded geospatial reasoning. VertiCue-Bench comprises 1,534 carefully curated instances across 17 tasks, explicitly disentangling low-level height perception from ambiguity-aware semantic reasoning. Evaluations on 14 state-of-the-art general and remote-sensing-specialized MLLMs, combined with counterfactual modality testing, reveal a striking perception-reasoning dissociation. While models exhibit emerging competence in reading raw CHM height cues, they largely fail to translate geometric perception into reliable semantic reasoning, often underperforming RGB-only baselines when joint constraints are required. Overall, VertiCue-Bench exposes a critical geometry-to-semantics gap in natural scene understanding, offering actionable insights for advancing geospatial MLLMs.

遥感大模型几何推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。