arXiv:2605.15024cs.CV2026-05

提出分层语义解耦框架,解决遥感图像变化描述中的语义混淆问题。

HiSem: Hierarchical Semantic Disentangling for Remote Sensing Image Change Captioning

论文配图:HiSem: Hierarchical Semantic Disentangling for Remote Sensing Image Change Captioning
图 1 · 摘自论文原文
  • 分层解耦设计:图像级粗粒度区分变与不变,标记级细粒度建模复杂变化
  • 在WHU-CDC数据集上提升7.52% BLEU-4,显著优于现有方法
  • 适合关注遥感变化理解、多时相图像分析的研究者

遥感图像变化描述(RSICC)旨在实现对双时相图像间真实变化的高层语义理解。尽管取得进展,现有方法受限于统一建模假设:变化与未变化图像对具有固有不同语义粒度,却采用相同处理策略,导致粗粒度变化存在判断与细粒度语义理解之间的语义纠缠。为此,我们提出一种新型分层语义解耦网络(HiSem),显式解耦不同粒度的语义表示。首先引入双向差异注意力调制(BDAM)模块,利用差异感知注意力增强跨时相交互,放大真实变化信号并抑制无关变化。在此基础上,设计分层自适应语义解耦(HASD)模块,在两级进行自适应路由:图像级粗粒度路由区分变化与未变化图像对,标记级细粒度混合专家(MoE)块建模变化样本中多样且异构的变化语义。在两个基准数据集上的大量实验表明,HiSem优于先前方法,在WHU-CDC数据集上实现+7.52% BLEU-4的显著提升。更重要的是,本方法为RSICC提供了结构化视角,明确将模型设计与双时相场景的内在语义异质性对齐。代码将公开于https://github.com/Man-Wang-star/HiSem。

原文摘要 · Abstract (English)

Remote sensing image change captioning (RSICC) aims to achieve high-level semantic understanding of genuine changes occurring between bi-temporal images. Despite notable progress, existing methods are fundamentally limited by a shared modeling assumption: changed and unchanged image pairs, which have intrinsically different semantic granularities, are processed under a unified modeling strategy. This modeling inconsistency leads to semantic entanglement between coarse-grained change existence judgment and fine-grained semantic understanding.To address the above limitation, we propose a novel hierarchical semantic disentangling network (HiSem) that explicitly disentangles semantic representations of different granularities. Specifically, we first introduce the Bidirectional Differential Attention Modulation (BDAM) module that leverages discrepancy-aware attention to enhance cross-temporal interactions, thereby amplifying true change signals while suppressing irrelevant variations. Building upon this, we design a Hierarchical Adaptive Semantic Disentanglement (HASD) module that performs adaptive routing at two hierarchical levels: a coarse-grained image-level routing mechanism distinguishes changed and unchanged image pairs, while a fine-grained token-level Mixture-of-Experts (MoE) block models diverse and heterogeneous change semantics for changed samples. Extensive experiments on two benchmark datasets demonstrate that HiSem outperfoms previous methods, achieving a significant improvement of +7.52\% BLEU-4 on the WHU-CDC dataset. More importantly, our approach provides a structured perspective for RSICC by explicitly aligning model design with the intrinsic semantic heterogeneity of bi-temporal scenes. The code will be available at https://github.com/Man-Wang-star/HiSem

遥感图像变化检测语义解耦视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。