跨模态遥感图像检索新模型,提升不同传感器间匹配准确率
CR-JEPA: Cross-Modal Joint-Embedding Predictive Learning for Remote Sensing Image Retrieval

- 采用双模态分支+共享主干的架构,通过预测目标特征实现跨模态对齐
- 在BEN-14K数据集上,S1到S2检索准确率提升至75.82%,超越基线14.59个百分点
- 设计解耦头结构,兼顾同模态与跨模态检索,参数更少且性能更强
跨模态遥感图像检索旨在跨异质传感模态检索语义相关场景。该任务极具挑战性,因配对观测在成像物理、空间分辨率、光谱配置和视觉外观上可能存在显著差异。此外,单一检索投影仅用一个目标函数难以同时支持跨模态语义对齐与同模态邻域保持。本文提出CR-JEPA,一种用于双模态遥感检索的交叉模态检索联合嵌入预测架构。模型采用模态特定的茎干、共享的Transformer主干以及类JEPA的预测目标,以估计模态内与模态间的掩码潜在目标特征。受LeJEPA启发,我们对原始检索投影应用草图各向同性高斯正则化以稳定嵌入并缓解崩溃问题。CR-JEPA进一步采用解耦头设计:统一检索头用于同模态检索,跨模态检索头用于跨模态搜索。我们在BEN-14K、CBRSIR_VS和DSRSID数据集上进行评估。在BEN-14K上,相较于X-JEPA,CR-JEPA将S1到S2检索准确率从61.23%提升至75.82%,S2到S1从63.73%提升至75.40%,同时以更少参数实现具有竞争力的同模态检索性能。
原文摘要 · Abstract (English)
Cross-modal remote sensing image retrieval aims to retrieve semantically related scenes across heterogeneous sensing modalities. This remains challenging because paired observations may differ substantially in imaging physics, spatial resolution, spectral configuration, and visual appearance. Moreover, a single retrieval projection trained with one objective may be insufficient to jointly support cross-modal semantic alignment and same-modal neighbourhood preservation. We propose CR-JEPA, a Cross-modal Retrieval Joint-Embedding Predictive Architecture for dual-modality remote sensing retrieval. The model uses modality-specific stems, a shared transformer trunk, and JEPA-style predictive objectives to estimate masked latent target features within and across modalities. Inspired by LeJEPA, we apply Sketched Isotropic Gaussian Regularization to raw retrieval projections to stabilize embeddings and mitigate collapse. CR-JEPA further employs a decoupled-head design with a unified retrieval head for same-modal retrieval and a cross-modal retrieval head for cross-modal search. We evaluate CR-JEPA on BEN-14K, CBRSIR_VS, and DSRSID. On BEN-14K, CR-JEPA improves S1 to S2 retrieval from 61.23% to 75.82% and S2 to S1 retrieval from 63.73% to 75.40% over X-JEPA, while also achieving competitive same-modal retrieval with fewer parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。