为遥感模型补足垂直维度感知,提升灾害场景理解能力
GeoHeight-Bench: Towards Height-Aware Multimodal Reasoning in Remote Sensing
- 用视觉语言模型自动生成带高程标注的数据,解决标注数据稀缺问题
- 构建两个层级的评测基准,支持相对高度与地形综合推理任务
- 提出首个垂向感知遥感大模型基线,验证高程信息对模型性能的关键作用
当前地球观测中的大型多模态模型通常忽略关键的“垂直”维度,限制了其在复杂遥感几何和灾害场景下的推理能力,而这些场景中物理空间结构往往比平面视觉纹理更重要。为此,我们提出一个全面的评估框架,专用于高程感知的遥感理解。首先,为克服标注数据严重匮乏的问题,我们开发了一种可扩展的、基于视觉语言模型的数据生成流程,结合系统化提示工程与元数据提取,构建了两个互补的基准:用于相对高程分析的GeoHeight-Bench,以及更具挑战性的、需全局地形感知的GeoHeight-Bench+。此外,为验证高程感知的必要性,我们提出了GeoHeightChat,首个面向遥感的高程感知大模型基线。该基线证明,将视觉语义与隐式注入的高程几何特征融合,能有效缓解“垂直盲区”,成功开启现有光学模型中交互式高程推理的新范式。
原文摘要 · Abstract (English)
Current Large Multimodal Models (LMMs) in Earth Observation typically neglect the critical "vertical" dimension, limiting their reasoning capabilities in complex remote sensing geometries and disaster scenarios where physical spatial structures often outweigh planar visual textures. To bridge this gap, we introduce a comprehensive evaluation framework dedicated to height-aware remote sensing understanding. First, to overcome the severe scarcity of annotated data, we develop a scalable, VLM-driven data generation pipeline utilizing systematic prompt engineering and metadata extraction. This pipeline constructs two complementary benchmarks: GeoHeight-Bench for relative height analysis, and a more challenging GeoHeight-Bench+ for holistic, terrain-aware reasoning. Furthermore, to validate the necessity of height perception, we propose GeoHeightChat, the first height-aware remote sensing LMM baseline. Serving as a strong proof of concept, our baseline demonstrates that synergizing visual semantics with implicitly injected height geometric features effectively mitigates the "vertical blind spot", successfully unlocking a new paradigm of interactive height reasoning in existing optical models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。