arXiv:2410.06194cs.CV2024-10被引 6

用文本提示让模型精准提取遥感图像中的建筑、道路等语义轮廓。

Prompting DirectSAM for Semantic Contour Extraction in Remote Sensing Images

  • 基于DirectSAM构建遥感专用模型,加入文本编码器与交叉注意力模块。
  • 在超3.4万组遥感图文轮廓数据上训练,性能超越现有方法。
  • 零样本和微调均表现优异,适合遥感语义分割场景应用。

Direct Segment Anything Model(DirectSAM)在无类别依赖的轮廓提取方面表现出色。本文探索其在光学遥感图像中的应用,此类图像中对建筑、道路网和海岸线等语义轮廓的提取具有重要实际价值。当前这些任务通常依赖于在小规模数据集上分别训练特定的小型模型。我们提出一个基于DirectSAM的遥感领域基础模型——DirectSAM-RS,不仅继承了自然图像中学习到的强大分割能力,还利用我们构建的大规模遥感语义轮廓数据集进行训练。该数据集包含超过34,000个图像-文本-轮廓三元组,比单个同类数据集大至少30倍。DirectSAM-RS集成了一个提示模块:包含文本编码器与交叉注意力层,可灵活地以目标类别标签或指代表达作为条件。我们在零样本和微调设置下评估了DirectSAM-RS,在多个下游基准测试中均达到了最先进的性能。

原文摘要 · Abstract (English)

The Direct Segment Anything Model (DirectSAM) excels in class-agnostic contour extraction. In this paper, we explore its use by applying it to optical remote sensing imagery, where semantic contour extraction-such as identifying buildings, road networks, and coastlines-holds significant practical value. Those applications are currently handled via training specialized small models separately on small datasets in each domain. We introduce a foundation model derived from DirectSAM, termed DirectSAM-RS, which not only inherits the strong segmentation capability acquired from natural images, but also benefits from a large-scale dataset we created for remote sensing semantic contour extraction. This dataset comprises over 34k image-text-contour triplets, making it at least 30 times larger than individual dataset. DirectSAM-RS integrates a prompter module: a text encoder and cross-attention layers attached to the DirectSAM architecture, which allows flexible conditioning on target class labels or referring expressions. We evaluate the DirectSAM-RS in both zero-shot and fine-tuning setting, and demonstrate that it achieves state-of-the-art performance across several downstream benchmarks.

遥感分割文本提示轮廓提取基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。