REJEPA通过联合嵌入预测提升遥感图像检索效率与精度。
REJEPA: A Novel Joint-Embedding Predictive Architecture for Efficient Remote Sensing Image Retrieval
- 用空间分布上下文编码预测目标特征,避免像素级细节干扰。
- 相比重建类模型计算量降低40%-60%,多数据集检索准确率提升5.1%-10.1%。
- 适用于多传感器、高密度复杂场景,适合高效遥感检索应用。
遥感图像库的快速膨胀亟需高效的内容检索技术。本文提出REJEPA(基于联合嵌入预测架构的检索),一种面向单模态遥感图像内容检索(RS-CBIR)的自监督框架。REJEPA采用空间分布上下文标记编码,预测目标标记的抽象表示,有效捕捉高层语义特征并去除冗余像素级细节。与关注像素重建的生成方法或依赖负样本对的对比学习不同,REJEPA在特征空间内运作,相较像素重建基线(如Masked Autoencoders, MAE)计算复杂度降低40%-60%。为保证表征强且多样,引入方差-不变性-协方差正则化(VICReg),防止编码器坍缩,减少特征冗余。在广泛使用的遥感基准BEN-14K(多光谱与SAR数据)、FMoW-RGB和FMoW-Sentinel上,相对于CSMAE-SESD、Mask-VLM、SatMAE、ScaleMAE、SatMAE++等先进自监督方法,平均检索准确率分别提升5.1%(BEN-14K S1)、7.4%(BEN-14K S2)、6.0%(FMoW-RGB)和10.1%(FMoW-Sentinel)。REJEPA在多种传感器模态间表现出良好泛化能力,确立了高效、可扩展、精准的遥感图像检索新基准,有效应对分辨率差异、高物体密度及复杂背景等挑战。
原文摘要 · Abstract (English)
The rapid expansion of remote sensing image archives demands the development of strong and efficient techniques for content-based image retrieval (RS-CBIR). This paper presents REJEPA (Retrieval with Joint-Embedding Predictive Architecture), an innovative self-supervised framework designed for unimodal RS-CBIR. REJEPA utilises spatially distributed context token encoding to forecast abstract representations of target tokens, effectively capturing high-level semantic features and eliminating unnecessary pixel-level details. In contrast to generative methods that focus on pixel reconstruction or contrastive techniques that depend on negative pairs, REJEPA functions within feature space, achieving a reduction in computational complexity of 40-60% when compared to pixel-reconstruction baselines like Masked Autoencoders (MAE). To guarantee strong and varied representations, REJEPA incorporates Variance-Invariance-Covariance Regularisation (VICReg), which prevents encoder collapse by promoting feature diversity and reducing redundancy. The method demonstrates an estimated enhancement in retrieval accuracy of 5.1% on BEN-14K (S1), 7.4% on BEN-14K (S2), 6.0% on FMoW-RGB, and 10.1% on FMoW-Sentinel compared to prominent SSL techniques, including CSMAE-SESD, Mask-VLM, SatMAE, ScaleMAE, and SatMAE++, on extensive RS benchmarks BEN-14K (multispectral and SAR data), FMoW-RGB and FMoW-Sentinel. Through effective generalisation across sensor modalities, REJEPA establishes itself as a sensor-agnostic benchmark for efficient, scalable, and precise RS-CBIR, addressing challenges like varying resolutions, high object density, and complex backgrounds with computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。