融合全局与局部信息,提升遥感图文检索精度
Cross-Modal Pre-Aligned Method with Global and Local Information for Remote-Sensing Image and Text Retrieval
- 设计双注意力结构捕捉遥感图像多尺度特征
- 预对齐机制使跨模态融合更高效,提升检索准确率
- 适合遥感图像与文本匹配、地理信息检索场景
遥感跨模态图文检索(RSCTIR)在信息挖掘中具有重要价值,但受限于遥感图像的多样性以及模态融合前缺乏有效特征对齐,导致检索精度与效率不足。为此,本文提出CMPAGL方法,通过全局与局部信息融合提升性能。其Gswin Transformer块结合局部窗口自注意力与全局-局部交叉注意力,实现多尺度特征提取;引入预对齐机制简化模态融合训练流程;设计相似性矩阵重加权(SMR)算法进行重排序,并在三元组损失中加入类内距离项优化特征学习。在四组数据集(包括RSICD和RSITMD)上的实验表明,该方法相较当前最优模型,在R@1上最高提升4.65%,平均召回率(mR)提升2.28%。
原文摘要 · Abstract (English)
Remote sensing cross-modal text-image retrieval (RSCTIR) has gained attention for its utility in information mining. However, challenges remain in effectively integrating global and local information due to variations in remote sensing imagery and ensuring proper feature pre-alignment before modal fusion, which affects retrieval accuracy and efficiency. To address these issues, we propose CMPAGL, a cross-modal pre-aligned method leveraging global and local information. Our Gswin transformer block combines local window self-attention and global-local window cross-attention to capture multi-scale features. A pre-alignment mechanism simplifies modal fusion training, improving retrieval performance. Additionally, we introduce a similarity matrix reweighting (SMR) algorithm for reranking, and enhance the triplet loss function with an intra-class distance term to optimize feature learning. Experiments on four datasets, including RSICD and RSITMD, validate CMPAGL's effectiveness, achieving up to 4.65% improvement in R@1 and 2.28% in mean Recall (mR) over state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。