arXiv:2604.26774cs.CVcs.AI2026-04被引 2

无需训练即可检测遥感图像语义变化,通过跨时序记忆推理提升准确率

MemOVCD: Training-Free Open-Vocabulary Change Detection via Cross-Temporal Memory Reasoning and Global-Local Adaptive Rectification

论文配图:MemOVCD: Training-Free Open-Vocabulary Change Detection via Cross-Temporal Memory Reasoning and Global-Local Adaptive Rectification
图 1 · 摘自论文原文
  • 将双时相变化检测视为两帧跟踪问题,双向传播聚合时间线索
  • 在五个基准上达到领先性能,显著提升对真实语义变化的识别能力
  • 适合遥感、环境监测等开放词汇场景,无需重新训练

开放词汇变化检测旨在不依赖预定义类别的情况下,识别双时相遥感图像中的语义变化。现有方法虽结合SAM、DINO和CLIP等基础模型,但通常独立处理每个时相或仅在最终对比阶段交互,导致语义推理中时序耦合不足,难以区分真实语义变化与非语义外观差异。此外,高分辨率图像上的局部块主导推理会削弱全局语义连续性,造成碎片化变化区域。为此,我们提出MemOVCD,一种基于跨时序记忆推理与全局-局部自适应修正的免训练开放词汇变化检测框架。具体而言,将双时相变化检测重构为两帧跟踪任务,引入加权双向传播以从两个时间方向聚合语义证据。为稳定长时距下的记忆传播,构建直方图对齐的过渡帧以平滑突变外观。此外,设计全局-局部自适应修正策略,动态融合局部与全局视图预测,在保持细粒度细节的同时提升空间一致性。在五个基准上的实验表明,MemOVCD在两项变化检测任务中均取得优异表现,验证了其在多样开放词汇设置下的有效性与泛化能力。

原文摘要 · Abstract (English)

Open-vocabulary change detection aims to identify semantic changes in bi-temporal remote sensing images without predefined categories. Recent methods combine foundation models such as SAM, DINO and CLIP, but typically process each timestamp independently or interact only at the final comparison stage. Such paradigms suffer from insufficient temporal coupling during semantic reasoning, which limits their ability to distinguish genuine semantic changes from non-semantic appearance discrepancies. In addition, patch-dominant inference on high-resolution images often weakens global semantic continuity and produces fragmented change regions. To address these issues, we propose MemOVCD, a training-free open-vocabulary change detection framework based on cross-temporal memory reasoning and global-local adaptive rectification. Specifically, we reformulate bi-temporal change detection as a two-frame tracking problem and introduce weighted bidirectional propagation to aggregate semantic evidence from both temporal directions. To stabilize memory propagation across large temporal gaps, we construct histogram-aligned transition frames to smooth abrupt appearance changes. Moreover, a global-local adaptive rectification strategy adaptively fuses local and global-view predictions, improving spatial consistency while preserving fine-grained details. Experiments on five benchmarks demonstrate that MemOVCD achieves favorable performance on two change detection tasks, validating its effectiveness and generalization under diverse open-vocabulary settings.

变化检测遥感开放词汇免训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。