用神经记忆提升遥感变化检测的长程建模能力
Towards Remote Sensing Change Detection with Neural Memory
- 基于钛模型设计视觉主干,结合分段局部注意力与神经记忆
- 在LEVIR-CD上达84.36%交并比和91.52%准确率
- 适合需要高效长程建模的遥感变化检测任务
遥感变化检测对环境监测和城市规划至关重要。现有方法难以在保持计算效率的同时捕捉长程依赖。尽管Transformer能建模全局上下文,但其二次复杂度带来可扩展性挑战;而现有线性注意力方法常无法捕捉复杂的时空关系。受语言任务中钛模型成功启发,我们提出ChangeTitans框架。具体地,提出VTitans,首个基于钛模型的视觉主干,融合神经记忆与分段局部注意力,有效捕获长程依赖并降低计算开销。进一步设计层级式VTitans-Adapter,优化多尺度特征。最后引入双流融合模块TS-CBAM,利用跨时序注意力抑制伪变化,提升检测精度。在四个基准数据集(LEVIR-CD、WHU-CD、LEVIR-CD+、SYSU-CD)上的实验表明,ChangeTitans达到当前最优性能,在LEVIR-CD上实现84.36%交并比和91.52%F1分数,同时保持计算竞争力。
原文摘要 · Abstract (English)
Remote sensing change detection is essential for environmental monitoring, urban planning, and related applications. However, current methods often struggle to capture long-range dependencies while maintaining computational efficiency. Although Transformers can effectively model global context, their quadratic complexity poses scalability challenges, and existing linear attention approaches frequently fail to capture intricate spatiotemporal relationships. Drawing inspiration from the recent success of Titans in language tasks, we present ChangeTitans, the Titans-based framework for remote sensing change detection. Specifically, we propose VTitans, the first Titans-based vision backbone that integrates neural memory with segmented local attention, thereby capturing long-range dependencies while mitigating computational overhead. Next, we present a hierarchical VTitans-Adapter to refine multi-scale features across different network layers. Finally, we introduce TS-CBAM, a two-stream fusion module leveraging cross-temporal attention to suppress pseudo-changes and enhance detection accuracy. Experimental evaluations on four benchmark datasets (LEVIR-CD, WHU-CD, LEVIR-CD+, and SYSU-CD) demonstrate that ChangeTitans achieves state-of-the-art results, attaining \textbf{84.36\%} IoU and \textbf{91.52\%} F1-score on LEVIR-CD, while remaining computationally competitive.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。