用扩散模型重写灾情变化描述,避免早期错误导致的连锁失误
EchoChange: A Diffusion Language Model with Dual Pass Remasking for Factual Remote Sensing Disaster Change Captioning

- 采用迭代去噪机制替代逐词生成,逐步修正描述内容
- 在RSCC数据集上各项指标超越现有方法,尤其提升语义准确性
- 适合需要高精度灾情文本生成的研究与应急响应应用
双时相遥感灾情变化描述需在大规模灾前灾后影像中识别稀疏且空间局部的变化,并转化为连贯准确的描述。然而现有方法多采用自回归解码生成文本,一旦早期误判变化对象、事件或空间关系,将不可逆地影响后续生成,加剧视觉模糊带来的事实性错误。为此,我们提出EchoChange,一种基于图像对条件的多模态离散扩散语言模型,将变化描述任务建模为迭代掩码词去噪过程,而非顺序生成。通过反复修正整个描述并结合图像对信息,模型可重新审视不确定内容,修正中间预测缺陷。进一步引入草稿感知双阶段训练、渐进式掩码课程及置信度引导重掩码策略,使训练与迭代推理对齐。在RSCC基准上的大量实验表明,EchoChange在词汇与语义指标上显著优于通用及遥感专用基线模型。
原文摘要 · Abstract (English)
Bi-temporal remote-sensing disaster change captioning often needs to identify sparse and spatially localized changes across large pre- and post-event scenes and then translate them into coherent, factual descriptions. However, existing change captioning methods always follow an autoregressive decoding paradigm to generate the change description and thus an early misinterpretation of the changed object, event, or spatial relation becomes an irreversible premise for subsequent text, amplifying visual ambiguity into cascading factual errors. To address this limitation, we propose EchoChange, a multimodal discrete diffusion language model that formulates change captioning as iterative masked-token denoising rather than left-to-right generation. By repeatedly revising the entire caption while conditioning on the image pair, EchoChange can reconsider uncertain content and correct imperfect intermediate predictions. We further introduce draft-aware dual-pass training, a progressive masking curriculum, and confidence-guided remasking to align training with iterative inference. Extensive experiments on the RSCC benchmark show that EchoChange substantially outperforms both general-purpose and remote-sensing-specific baselines across lexical and semantic metrics. The EchoChange Project is at https://sundongwei.github.io/EchoChange_Project/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。