arXiv:2503.07266cs.CV2025-03AAAI被引 5

让SAM2精准理解遥感图像描述,实现更准的语义分割。

RS2-SAM2: Customized SAM2 for Referring Remote Sensing Image Segmentation

  • 用联合编码器对齐遥感图像与文本特征
  • 生成伪掩码作为密集提示,提升定位精度
  • 适合需要精准语义分割的遥感应用

指代式遥感图像分割(RRSIS)旨在根据文本描述分割遥感图像中的目标物体。尽管分割一切模型2(SAM2)在多种分割任务中表现优异,但其应用于RRSIS仍面临挑战,包括理解文本描述的遥感场景以及从文本生成有效提示。为此,我们提出RS2-SAM2,一种通过对齐适配后的遥感特征与文本特征,并提供基于伪掩码的密集提示,来适配SAM2至RRSIS的新框架。具体而言,我们采用联合编码器联合编码视觉与文本输入,生成对齐的视觉与文本嵌入及多模态类别标记;引入双向层级融合模块,适应遥感场景并对齐适配后的视觉特征与视觉增强的文本嵌入,提升模型对文本描述遥感场景的理解能力。为向SAM2提供精确的目标线索,设计掩码提示生成器,以视觉嵌入和类别标记为输入,生成伪掩码作为SAM2的密集提示。多个RRSIS基准测试结果表明,RS2-SAM2达到当前最优性能。

原文摘要 · Abstract (English)

Referring Remote Sensing Image Segmentation (RRSIS) aims to segment target objects in remote sensing (RS) images based on textual descriptions. Although Segment Anything Model 2 (SAM2) has shown remarkable performance in various segmentation tasks, its application to RRSIS presents several challenges, including understanding the text-described RS scenes and generating effective prompts from text. To address these issues, we propose \textbf{RS2-SAM2}, a novel framework that adapts SAM2 to RRSIS by aligning the adapted RS features and textual features while providing pseudo-mask-based dense prompts. Specifically, we employ a union encoder to jointly encode the visual and textual inputs, generating aligned visual and text embeddings as well as multimodal class tokens. A bidirectional hierarchical fusion module is introduced to adapt SAM2 to RS scenes and align adapted visual features with the visually enhanced text embeddings, improving the model's interpretation of text-described RS scenes. To provide precise target cues for SAM2, we design a mask prompt generator, which takes the visual embeddings and class tokens as input and produces a pseudo-mask as the dense prompt of SAM2. Experimental results on several RRSIS benchmarks demonstrate that RS2-SAM2 achieves state-of-the-art performance.

遥感分割多模态SAM2提示生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。