用新型视觉架构提升遥感变化检测精度与效率
NeXt2Former-CD: Efficient Remote Sensing Change Detection with Modern Vision Architectures
- 融合卷积与注意力机制,增强对配准误差和小目标偏移的容忍度
- 在三个数据集上F1与IoU均超越现有Mamba基线方法
- 参数量更大但推理速度相近,适合高分辨率应用
状态空间模型(SSMs)因其良好的扩展性在遥感变化检测中备受关注。本文探索现代卷积与注意力架构作为替代方案的潜力,提出NeXt2Former-CD端到端框架:采用以DINOv3权重初始化的双边卷积神经网络编码器、基于可变形注意力的时序融合模块及Mask2Former解码器。该设计旨在更好应对双时相影像中的残余配准噪声、微小空间偏移及语义模糊问题。在LEVIR-CD、WHU-CD和CDD数据集上的实验表明,该方法在所有评估方法中表现最佳,其F1分数和交并比(IoU)均优于近期基于Mamba的基线。尽管参数量更大,模型推理延迟仍与SSM方法相当,表明其在高分辨率变化检测任务中具备实际可行性。
原文摘要 · Abstract (English)
State Space Models (SSMs) have recently gained traction in remote sensing change detection (CD) for their favorable scaling properties. In this paper, we explore the potential of modern convolutional and attention-based architectures as a competitive alternative. We propose NeXt2Former-CD, an end-to-end framework that integrates a Siamese ConvNeXt encoder initialized with DINOv3 weights, a deformable attention-based temporal fusion module, and a Mask2Former decoder. This design is intended to better tolerate residual co-registration noise and small object-level spatial shifts, as well as semantic ambiguity in bi-temporal imagery. Experiments on LEVIR-CD, WHU-CD, and CDD datasets show that our method achieves the best results among the evaluated methods, improving over recent Mamba-based baselines in both F1 score and IoU. Furthermore, despite a larger parameter count, our model maintains inference latency comparable to SSM-based approaches, suggesting it is practical for high-resolution change detection tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。