解决遥感图文检索持续学习中的跨模态失准问题。
DARAD: Dual Adapters and Ranking-Aware Distillation for Continual Remote Sensing Image-Text Retrieval

- 双适配器融合区域与像素级视觉特征,应对尺度变化。
- 文本分支用多专家路由分离通用与专用语义,减少漂移。
- 双向排名蒸馏保留历史检索结构,适合增量数据场景。
随着地球观测技术的快速发展,遥感档案持续膨胀,遥感图像-文本检索(RS-ITR)日益重要。然而,由于遥感数据在尺度和分布上的变化,导致跨模态对齐空间失真,现有持续学习(CL)方法难以支撑可靠持续检索。为此,我们提出DARAD框架,通过双适配器与排名感知蒸馏,既保留历史跨模态排序结构,又能从不断演化的数据中学习新视觉与文本概念。视觉分支引入空间融合适配器,整合粗粒度区域与细粒度块级线索,以适应遥感尺度变化,并将视觉更新锚定在预训练对齐空间。文本分支采用多专家语义路由,分离共享语义与语义特化残差,吸收新出现描述的同时限制全局文本嵌入漂移。此外,双向排名蒸馏利用冻结教师模型与历史锚点,保留历史跨模态排序结构,缓解持续阶段间的对齐空间失真。在多阶段持续检索协议下实验表明,DARAD显著优于现有CL方法,在适应新数据的同时保持对历史数据的有效性。
原文摘要 · Abstract (English)
With the rapid growth of Earth observation technologies, remote sensing archives are rapidly expanding, making remote sensing image-text retrieval (RS-ITR) increasingly important. However, continual RS-ITR remains challenging because scale variation and distribution shifts in RS aggravate cross-modal alignment space distortion, making it difficult for existing continual learning (CL) methods to support reliable continual retrieval. To address this challenge, we propose DARAD, a dual-adapter and ranking-aware distillation framework that preserves the historical cross-modal ranking structure while learning new visual and textual concepts from evolving archives. Specifically, the visual branch introduces a spatial fusion adapter, which integrates coarse regional cues and fine-grained patch cues to accommodate RS scale variation while anchoring visual updates to the pretrained alignment space. The textual branch employs multi-expert semantic routing, which separates shared textual semantics from semantically specialized residuals to absorb newly emerging descriptions while constraining global text embedding drift. Furthermore, bidirectional ranking distillation uses a frozen teacher model and historical anchors to preserve the historical cross-modal ranking structure, thereby mitigating alignment space distortion across continual stages. Experiments under a multi-stage continual retrieval protocol show that DARAD achieves superior performance over existing CL methods, improving adaptation to newly arrived data while maintaining effectiveness on historical data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。