arXiv:2504.09644cs.CV2025-04被引 51

用大模型理解地理图像中的隐含问题并生成目标区域掩码

SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model

  • 通过语言模型解析隐式查询,融合多尺度视觉特征进行推理
  • 在5434张图上实现3万+隐式问答,性能超越传统与现有大模型方法
  • 适合遥感、城市规划、灾害管理等需要空间语义理解的场景

遥感在环境动态监测、城市规划和灾害管理中至关重要,但传统工作流依赖显式分割或检测,难以处理需空间上下文、领域知识和隐含意图推理的复杂查询。为此,我们提出新任务——地理空间像素推理,支持隐式提问与推理,并生成目标区域掩码。为推动该任务发展,我们构建并发布了首个大规模基准数据集EarthReason,包含5,434个手工标注图像掩码及超过30,000个隐式问答对。同时提出SegEarth-R1,一种简单有效的语言引导分割基线:采用分层视觉编码器、用于指令解析的大语言模型(LLM)和定制掩码生成器以捕捉空间相关性。其设计包含领域特化改进:激进的视觉令牌压缩以处理超高清遥感图像,描述投影模块融合语言与多尺度特征,以及直接查询描述嵌入的精简掩码预测流程。大量实验表明,SegEarth-R1在推理与指代分割任务上均达到领先水平,显著优于传统与基于大模型的分割方法。数据与代码将开源。

原文摘要 · Abstract (English)

Remote sensing has become critical for understanding environmental dynamics, urban planning, and disaster management. However, traditional remote sensing workflows often rely on explicit segmentation or detection methods, which struggle to handle complex, implicit queries that require reasoning over spatial context, domain knowledge, and implicit user intent. Motivated by this, we introduce a new task, \ie, geospatial pixel reasoning, which allows implicit querying and reasoning and generates the mask of the target region. To advance this task, we construct and release the first large-scale benchmark dataset called EarthReason, which comprises 5,434 manually annotated image masks with over 30,000 implicit question-answer pairs. Moreover, we propose SegEarth-R1, a simple yet effective language-guided segmentation baseline that integrates a hierarchical visual encoder, a large language model (LLM) for instruction parsing, and a tailored mask generator for spatial correlation. The design of SegEarth-R1 incorporates domain-specific adaptations, including aggressive visual token compression to handle ultra-high-resolution remote sensing images, a description projection module to fuse language and multi-scale features, and a streamlined mask prediction pipeline that directly queries description embeddings. Extensive experiments demonstrate that SegEarth-R1 achieves state-of-the-art performance on both reasoning and referring segmentation tasks, significantly outperforming traditional and LLM-based segmentation methods. Our data and code will be released at https://github.com/earth-insights/SegEarth-R1.

遥感大模型图像分割空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。