arXiv:2505.16756cs.CVcs.IR2025-05被引 5

提出不对称适配器缓解遥感图文检索中模态失衡问题

Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval

  • 设计异构适配器,分别增强图像细节与文本语义提取
  • 双任务一致性损失提升跨模态对齐鲁棒性,mR提升6%-11%
  • 适合需要轻量化训练的遥感图文检索场景

遥感图像-文本检索(RSITR)在地理信息解析、灾害监测和城市规划中至关重要,通过建立图像与文本描述之间的语义关联发挥作用。现有视觉-语言预训练模型的参数高效微调(PEFT)方法通常采用对称适配器结构以探索跨模态相关性,但文本模态强区分性会主导优化过程,抑制图像表征学习,导致显著的跨模态优化不平衡,成为性能提升的瓶颈。为此,本文提出一种表示差异桥接(RDB)方法。一方面,设计跨模态非对称适配器(CMAA),包含视觉增强适配器(VEA)与文本语义适配器(TSA),分别通过差分注意力机制挖掘图像细粒度特征,通过分层注意力机制识别关键文本语义;另一方面,将传统单任务检索框架扩展为双任务优化框架,引入双任务一致性损失(DTCL),通过自适应加权组合跨模态、分类与指数移动平均一致性约束,提升对齐鲁棒性。在RSICD与RSITMD数据集上的实验表明,所提RDB方法相比先进PEFT方法在mR指标上提升6%-11%,优于全量微调的GeoRSCLIP模型1.15%-2%。

原文摘要 · Abstract (English)

Remote Sensing Image-Text Retrieval (RSITR) plays a critical role in geographic information interpretation, disaster monitoring, and urban planning by establishing semantic associations between image and textual descriptions. Existing Parameter-Efficient Fine-Tuning (PEFT) methods for Vision-and-Language Pre-training (VLP) models typically adopt symmetric adapter structures for exploring cross-modal correlations. However, the strong discriminative nature of text modality may dominate the optimization process and inhibits image representation learning. The nonnegligible imbalanced cross-modal optimization remains a bottleneck to enhancing the model performance. To address this issue, this study proposes a Representation Discrepancy Bridging (RDB) method for the RSITR task. On the one hand, a Cross-Modal Asymmetric Adapter (CMAA) is designed to enable modality-specific optimization and improve feature alignment. The CMAA comprises a Visual Enhancement Adapter (VEA) and a Text Semantic Adapter (TSA). VEA mines fine-grained image features by Differential Attention (DA) mechanism, while TSA identifies key textual semantics through Hierarchical Attention (HA) mechanism. On the other hand, this study extends the traditional single-task retrieval framework to a dual-task optimization framework and develops a Dual-Task Consistency Loss (DTCL). The DTCL improves cross-modal alignment robustness through an adaptive weighted combination of cross-modal, classification, and exponential moving average consistency constraints. Experiments on RSICD and RSITMD datasets show that the proposed RDB method achieves a 6%-11% improvement in mR metrics compared to state-of-the-art PEFT methods and a 1.15%-2% improvement over the full fine-tuned GeoRSCLIP model.

遥感检索跨模态对齐轻量化微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。