构建高分辨率遥感变化描述数据集,提升复杂场景下文本生成准确性
Robust Change Captioning in Remote Sensing: SECOND-CC Dataset and MModalCC Framework
- 融合视觉与语义信息,通过跨模态注意力机制对齐图像与文本
- 在6041对遥感图像上实现BLEU4+4.6%、CIDEr+9.6%的性能提升
- 适合遥感变化分析、多模态生成与灾害监测方向的研究者
遥感变化描述(RSICC)旨在用自然语言描述双时相图像间的差异。现有方法在光照变化、视角偏移、模糊效应等挑战下表现不佳,尤其在无变化区域易产生错误描述。此外,不同空间分辨率图像及配准误差也影响生成质量。为此,我们提出SECOND-CC数据集,包含6,041对高分辨率RGB遥感图像对、语义分割图和多样化真实场景,共30,205条描述性句子。同时,我们设计MModalCC框架,通过交叉模态交叉注意力(CMCA)和多模态门控交叉注意力(MGCA)融合视觉与语义信息。消融实验与注意力可视化验证其有效性。全面实验表明,MModalCC优于最新方法如RSICCformer、Chg2Cap和PSNet,BLEU4提升4.6%,CIDEr提升9.6%。代码与数据集将公开于https://github.com/ChangeCapsInRS/SecondCC。
原文摘要 · Abstract (English)
Remote sensing change captioning (RSICC) aims to describe changes between bitemporal images in natural language. Existing methods often fail under challenges like illumination differences, viewpoint changes, blur effects, leading to inaccuracies, especially in no-change regions. Moreover, the images acquired at different spatial resolutions and have registration errors tend to affect the captions. To address these issues, we introduce SECOND-CC, a novel RSICC dataset featuring high-resolution RGB image pairs, semantic segmentation maps, and diverse real-world scenarios. SECOND-CC which contains 6,041 pairs of bitemporal RS images and 30,205 sentences describing the differences between images. Additionally, we propose MModalCC, a multimodal framework that integrates semantic and visual data using advanced attention mechanisms, including Cross-Modal Cross Attention (CMCA) and Multimodal Gated Cross Attention (MGCA). Detailed ablation studies and attention visualizations further demonstrate its effectiveness and ability to address RSICC challenges. Comprehensive experiments show that MModalCC outperforms state-of-the-art RSICC methods, including RSICCformer, Chg2Cap, and PSNet with +4.6% improvement on BLEU4 score and +9.6% improvement on CIDEr score. We will make our dataset and codebase publicly available to facilitate future research at https://github.com/ChangeCapsInRS/SecondCC
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。