提出新方法提升遥感图像变化检测精度,尤其在数据少时表现更优。
SChanger: Change Detection from a Semantic Change and Spatial Consistency Perspective
- 用预训练+微调策略学习语义特征,聚焦变化区域。
- 引入空间一致性注意力机制,增强多尺度变化建模能力。
- 在6个数据集上超越现有方法,最高达97.62%的准确率。
变化检测是地球观测应用中的关键任务。近年来深度学习方法展现出强大性能和广泛应用。然而,由于精确配准同一地区遥感图像过程耗时费力,导致数据稀缺,限制了深度学习算法的表现。为解决数据稀缺问题,我们提出一种称为语义变化网络(SCN)的微调策略。首先在单时相监督任务上预训练模型,获取实例特征提取先验知识;随后采用共享权重的双流结构和扩展的时间融合模块(TFM),在变化检测任务上微调,将原模型对所有实例的语义识别转变为仅关注变化部分。同时观察到两幅图像中变化位置具有空间一致性,我们通过大卷积核生成注意力图,施加于两个时间点的特征,强化多尺度变化建模并捕捉深层语义关联。构建了一个二值变化检测模型,在六个数据集上验证,均优于现有基准方法,分别取得LEVIR-CD、LEVIR-CD+、S2Looking、CDD、SYSU-CD、WHU-CD上的F1分数92.87%、86.43%、68.95%、97.62%、84.58%、93.20%。
原文摘要 · Abstract (English)
Change detection is a key task in Earth observation applications. Recently, deep learning methods have demonstrated strong performance and widespread application. However, change detection faces data scarcity due to the labor-intensive process of accurately aligning remote sensing images of the same area, which limits the performance of deep learning algorithms. To address the data scarcity issue, we develop a fine-tuning strategy called the Semantic Change Network (SCN). We initially pre-train the model on single-temporal supervised tasks to acquire prior knowledge of instance feature extraction. The model then employs a shared-weight Siamese architecture and extended Temporal Fusion Module (TFM) to preserve this prior knowledge and is fine-tuned on change detection tasks. The learned semantics for identifying all instances is changed to focus on identifying only the changes. Meanwhile, we observe that the locations of changes between the two images are spatially identical, a concept we refer to as spatial consistency. We introduce this inductive bias through an attention map that is generated by large-kernel convolutions and applied to the features from both time points. This enhances the modeling of multi-scale changes and helps capture underlying relationships in change detection semantics. We develop a binary change detection model utilizing these two strategies. The model is validated against state-of-the-art methods on six datasets, surpassing all benchmark methods and achieving F1 scores of 92.87%, 86.43%, 68.95%, 97.62%, 84.58%, and 93.20% on the LEVIR-CD, LEVIR-CD+, S2Looking, CDD, SYSU-CD, and WHU-CD datasets, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。