提出首个覆盖多维度的遥感图像语言分割数据集与模型
SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing Images
- 构建涵盖层级粒度、目标数量等四维度的大型遥感语言分割数据集
- 在复杂指令下实现小目标定位与多目标分割,性能显著优于现有方法
- 适合遥感智能分析、灾害响应等需要复杂语义理解的场景
在遥感图像中准确将复杂语言指令映射到像素,对灾害响应和环境监测至关重要。当前模型仅能处理简单单目标命令,在面对多尺度、多目标或隐含意图等复杂地理空间场景时表现不佳。为此,我们提出了LaSeRS——首个大规模数据集,覆盖四个关键维度:层级粒度、目标多重性、推理需求和语言多样性,突破传统数据集的简化局限,为复杂地理空间推理提供基准。同时,我们提出SegEarth-R2,一种面向遥感图像全面语言引导分割的多模态大模型架构。该模型通过空间注意力监督机制有效定位小目标及其部件,并采用灵活高效的分割查询机制,支持单目标与多目标场景。实验表明,SegEarth-R2在LaSeRS及其他基准上均取得优异性能,为下一代地理空间分割建立强大基线。所有数据与代码将公开于https://github.com/earth-insights/SegEarth-R2。
原文摘要 · Abstract (English)
Effectively grounding complex language to pixels in remote sensing (RS) images is a critical challenge for applications like disaster response and environmental monitoring. Current models can parse simple, single-target commands but fail when presented with complex geospatial scenarios, e.g., segmenting objects at various granularities, executing multi-target instructions, and interpreting implicit user intent. To drive progress against these failures, we present LaSeRS, the first large-scale dataset built for comprehensive training and evaluation across four critical dimensions of language-guided segmentation: hierarchical granularity, target multiplicity, reasoning requirements, and linguistic variability. By capturing these dimensions, LaSeRS moves beyond simple commands, providing a benchmark for complex geospatial reasoning. This addresses a critical gap: existing datasets oversimplify, leading to sensitivity-prone real-world models. We also propose SegEarth-R2, an MLLM architecture designed for comprehensive language-guided segmentation in RS, which directly confronts these challenges. The model's effectiveness stems from two key improvements: (1) a spatial attention supervision mechanism specifically handles the localization of small objects and their components, and (2) a flexible and efficient segmentation query mechanism that handles both single-target and multi-target scenarios. Experimental results demonstrate that our SegEarth-R2 achieves outstanding performance on LaSeRS and other benchmarks, establishing a powerful baseline for the next generation of geospatial segmentation. All data and code will be released at https://github.com/earth-insights/SegEarth-R2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。