arXiv:2604.20429cs.CV2026-04

提出分阶段检索框架,提升遥感图文匹配效率与精度。

Fast-then-Fine: A Two-Stage Framework with Multi-Granular Representation for Cross-Modal Retrieval in Remote Sensing

论文配图:Fast-then-Fine: A Two-Stage Framework with Multi-Granular Representation for Cross-Modal Retrieval in Remote Sensing
图 1 · 摘自论文原文
  • 先快速粗粒度召回候选,再精细文本引导重排序
  • 无需额外参数,实现跨模态细粒度对齐
  • 适合大规模遥感数据高效检索场景

遥感图像-文本检索在理解海量遥感影像中起关键作用。然而,遥感图像中密集多目标分布和复杂背景使得同时实现细粒度跨模态对齐与高效检索极具挑战。现有方法或依赖复杂的跨模态交互导致检索效率低,或依赖大规模视觉语言模型预训练,需大量数据与计算资源。为此,我们提出一种快-精(Fast-then-Fine, FTF)两阶段检索框架,将检索分解为无文本依赖的召回阶段和文本引导的重排序阶段。召回阶段采用无文本依赖的粗粒度表示进行高效候选选择;重排序阶段引入无参数的平衡文本引导交互模块,实现细粒度对齐且不增加可学习参数。此外,设计了跨模态损失函数,联合优化多粒度表示下的跨模态对齐。在公开基准上的大量实验表明,FTF在保持竞争力检索精度的同时,显著提升了检索效率。

原文摘要 · Abstract (English)

Remote sensing (RS) image-text retrieval plays a critical role in understanding massive RS imagery. However, the dense multi-object distribution and complex backgrounds in RS imagery make it difficult to simultaneously achieve fine-grained cross-modal alignment and efficient retrieval. Existing methods either rely on complex cross-modal interactions that lead to low retrieval efficiency, or depend on large-scale vision-language model pre-training, which requires massive data and computational resources. To address these issues, we propose a fast-then-fine (FTF) two-stage retrieval framework that decomposes retrieval into a text-agnostic recall stage for efficient candidate selection and a text-guided rerank stage for fine-grained alignment. Specifically, in the recall stage, text-agnostic coarse-grained representations are employed for efficient candidate selection; in the rerank stage, a parameter-free balanced text-guided interaction block enhances fine-grained alignment without introducing additional learnable parameters. Furthermore, an inter- and intra-modal loss is designed to jointly optimize cross-modal alignment across multi-granular representations. Extensive experiments on public benchmarks demonstrate that the FTF achieves competitive retrieval accuracy while significantly improving retrieval efficiency compared with existing methods.

遥感检索两阶段框架跨模态对齐高效检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。