arXiv:2411.06091cs.CV2024-11被引 13

针对遥感图像中多物体场景,提出自动聚合相似模式的自监督学习框架。

Pattern Integration and Enhancement Vision Transformer for Self-Supervised Learning in Remote Sensing

  • 设计教师-学生架构,结合地理模式聚类与特征融合重建。
  • 在多个下游任务中显著提升检测、分割和变化检测性能。
  • 适合遥感图像理解中缺乏明确前景目标的场景应用。

近期自监督学习方法在无标签遥感图像上取得了显著进展,但多数遥感图像以包含多个地物的场景为主,缺乏明确前景目标,限制了现有方法的性能。本文提出一种专为遥感影像设计的新型自监督学习框架——模式整合增强视觉变换器(PIEViT)。该框架采用教师-学生架构,同时处理图像级与块级任务。引入地理模式一致性(GPC)模块,探索图像块的自然聚类特性,增强个体特征的区分度;并通过特征整合投影(FIP)模块,利用空间聚类后的块进行掩码标记重建优化。在多个下游任务(包括目标检测、语义分割、变化检测)上验证了PIEViT的有效性,结果表明其显著提升了内部块特征表示能力,优于现有自监督基线,在目标检测、土地覆盖分类和变化检测任务中均表现优异,展现出强大的鲁棒性、泛化性与可迁移性。

原文摘要 · Abstract (English)

Recent self-supervised learning (SSL) methods have demonstrated impressive results in learning visual representations from unlabeled remote sensing images. However, most remote sensing images predominantly consist of scenographic scenes containing multiple ground objects without explicit foreground targets, which limits the performance of existing SSL methods that focus on foreground targets. This raises the question: Is there a method that can automatically aggregate similar objects within scenographic remote sensing images, thereby enabling models to differentiate knowledge embedded in various geospatial patterns for improved feature representation? In this work, we present the Pattern Integration and Enhancement Vision Transformer (PIEViT), a novel self-supervised learning framework designed specifically for remote sensing imagery. PIEViT utilizes a teacher-student architecture to address both image-level and patch-level tasks. It employs the Geospatial Pattern Cohesion (GPC) module to explore the natural clustering of patches, enhancing the differentiation of individual features. The Feature Integration Projection (FIP) module further refines masked token reconstruction using geospatially clustered patches. We validated PIEViT across multiple downstream tasks, including object detection, semantic segmentation, and change detection. Experiments demonstrated that PIEViT enhances the representation of internal patch features, providing significant improvements over existing self-supervised baselines. It achieves excellent results in object detection, land cover classification, and change detection, underscoring its robustness, generalization, and transferability for remote sensing image interpretation tasks.

自监督学习遥感图像视觉变换器模式聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。