arXiv:2503.12404cs.CV2025-03被引 5

用SAM2改进遥感图像分割标注,低成本生成高质量标签。

SAM2-ELNet: Label Enhancement and Automatic Annotation for Remote Sensing Segmentation

  • 基于SAM2冻结主干,微调适配器与解码器实现自动标注。
  • 仅用30%数据训练即生成可靠标签,性能略降但成本大减。
  • 适合资源有限的遥感任务,如海洋漏油、高分辨率道路分割。

遥感图像分割对环境监测、灾害评估和资源管理至关重要,但其性能高度依赖数据集质量。尽管已有多个高质量数据集,针对特定任务(如海洋漏油分割)仍面临数据稀缺问题,目前仍依赖耗时且主观的人工标注。虽然段落任意模型2(SAM2)具备自动标注潜力,但在异质性高、对比度低的遥感影像上表现不佳。为此,我们提出新型标签增强与自动标注框架SAM2-ELNet(Enhancement and Labeling Network)。具体地,采用预训练SAM2中的冻结Hiera主干作为编码器,微调适配器与解码器以适应不同遥感任务。此外,框架包含标签质量评估模块用于过滤,保障生成标签可靠性。我们在两个数据集上开展实验:使用合成孔径雷达(SAR)影像的Deep-SAR Oil Spill(SOS)数据集,以及使用超高清光学影像的CHN6-CUG Road数据集。该框架可增强粗略标注,并在资源受限条件下生成可靠训练数据。仅需30%原始训练数据微调,即可自动生成标注数据;仅使用这些自动生成数据训练的模型,性能略低于全量人工标注,但大幅降低标注成本,为大规模遥感解析提供实用方案。

原文摘要 · Abstract (English)

Remote sensing image segmentation is crucial for environmental monitoring, disaster assessment, and resource management, but its performance largely depends on the quality of the dataset. Although several high-quality datasets are broadly accessible, data scarcity remains for specialized tasks like marine oil spill segmentation. Such tasks still rely on manual annotation, which is both time-consuming and influenced by subjective human factors. The segment anything model 2 (SAM2) has strong potential as an automatic annotation framework but struggles to perform effectively on heterogeneous, low-contrast remote sensing imagery. To address these challenges, we introduce a novel label enhancement and automatic annotation framework, termed SAM2-ELNet (Enhancement and Labeling Network). Specifically, we employ the frozen Hiera backbone from the pretrained SAM2 as the encoder, while fine-tuning the adapter and decoder for different remote sensing tasks. In addition, the proposed framework includes a label quality evaluator for filtering, ensuring the reliability of the generated labels. We design a series of experiments targeting resource-limited remote sensing tasks and evaluate our method on two datasets: the Deep-SAR Oil Spill (SOS) dataset with Synthetic Aperture Radar (SAR) imagery, and the CHN6-CUG Road dataset with Very High Resolution (VHR) optical imagery. The proposed framework can enhance coarse annotations and generate reliable training data under resource-limited conditions. Fine-tuned on only 30% of the training data, it generates automatically labeled data. A model trained solely on these achieves slightly lower performance than using the full original annotations, while greatly reducing labeling costs and offering a practical solution for large-scale remote sensing interpretation.

遥感分割自动标注SAM2标签增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。