用SAM提升遥感图像开放词汇分割,无需标注新类别。
AerOSeg: Harnessing SAM for Open-Vocabulary Segmentation in Remote Sensing Images
- 用多旋转图像+领域提示提取稳定图文特征。
- 结合SAM特征,提升分割精度,平均提升2.54 h-mIoU。
- 适合遥感中未知类别识别,降低标注依赖。
遥感图像中超越预设类别的图像分割是关键挑战,因推理时常出现新类别。开放词汇分割可解决传统监督模型泛化不足问题,减少对昂贵像素级标注的依赖。现有开放词汇分割方法多针对自然图像,难以应对遥感数据中的尺度变化、方向差异与复杂场景。为此,本文提出AerOSeg,一种专为遥感设计的开放词汇分割方法。通过输入图像的多旋转版本与领域特定提示计算鲁棒的图文相关特征,并经空间与类别细化模块优化。借鉴分割一切模型(SAM)在多领域的成功,利用其特征引导相关特征的空间细化。同时引入语义反投影模块与损失,确保SAM语义信息在分割流程中无缝传递。最后通过多尺度注意力感知解码器增强特征,生成最终分割图。在iSAID、DLRSD和OpenEarthMap三个基准遥感数据集上验证,AerOSeg显著优于当前最优开放词汇分割方法,平均提升2.54 h-mIoU。
原文摘要 · Abstract (English)
Image segmentation beyond predefined categories is a key challenge in remote sensing, where novel and unseen classes often emerge during inference. Open-vocabulary image Segmentation addresses these generalization issues in traditional supervised segmentation models while reducing reliance on extensive per-pixel annotations, which are both expensive and labor-intensive to obtain. Most Open-Vocabulary Segmentation (OVS) methods are designed for natural images but struggle with remote sensing data due to scale variations, orientation changes, and complex scene compositions. This necessitates the development of OVS approaches specifically tailored for remote sensing. In this context, we propose AerOSeg, a novel OVS approach for remote sensing data. First, we compute robust image-text correlation features using multiple rotated versions of the input image and domain-specific prompts. These features are then refined through spatial and class refinement blocks. Inspired by the success of the Segment Anything Model (SAM) in diverse domains, we leverage SAM features to guide the spatial refinement of correlation features. Additionally, we introduce a semantic back-projection module and loss to ensure the seamless propagation of SAM's semantic information throughout the segmentation pipeline. Finally, we enhance the refined correlation features using a multi-scale attention-aware decoder to produce the final segmentation map. We validate our SAM-guided Open-Vocabulary Remote Sensing Segmentation model on three benchmark remote sensing datasets: iSAID, DLRSD, and OpenEarthMap. Our model outperforms state-of-the-art open-vocabulary segmentation methods, achieving an average improvement of 2.54 h-mIoU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。