arXiv:2605.31068cs.CV2026-05被引 1

用量子相似性增强遥感多模态表征学习,提升语义理解能力。

HQ-JEPA: Hybrid Quantum Joint-Embedding Predictive Architecture for Cross-Modal Remote Sensing Representation Learning

论文配图:HQ-JEPA: Hybrid Quantum Joint-Embedding Predictive Architecture for Cross-Modal Remote Sensing Representation Learning
图 1 · 摘自论文原文
  • 融合量子态重叠度设计新损失函数,优化跨模态特征对齐。
  • 在GeoBench上线性探测与微调均超越主流自监督模型。
  • 适合关注遥感多源数据融合与量子机器学习的科研人员。

我们提出HQ-JEPA,一种用于跨模态遥感表征学习的混合量子-经典联合嵌入预测架构。该框架将JEPA风格的掩码潜在表示预测扩展至配对的哨兵1号与哨兵2号影像,通过可见上下文区域预测被掩码的目标表示,并在共享嵌入空间中对齐异构模态特征。为提升表征质量,HQ-JEPA结合四项互补目标:潜在标记预测、跨模态标记对齐、基于SIGReg的融合潜在空间高斯正则化,以及基于可微分SWAP-test的保真度量子相似性(FQS)损失。与像素重建方法不同,HQ-JEPA直接在潜在空间学习语义表示,并引入基于量子态重叠的相似性作为额外正则信号。我们在GeoBench分类与分割任务上评估预训练编码器,在线性探测与微调设置下,结果表明其性能优于或相当强于现有自监督与遥感基础模型,验证了预测式自监督、跨模态几何正则化与量子保真度表征学习在遥感中的协同优势。

原文摘要 · Abstract (English)

We introduce HQ-JEPA, a hybrid quantum-classical joint-embedding predictive architecture for cross-modal remote sensing representation learning. The proposed framework extends JEPA-style masked latent prediction to paired Sentinel-1 and Sentinel-2 imagery by predicting masked target representations from visible context regions while aligning heterogeneous modality features in a shared embedding space. To improve representation quality, HQ-JEPA combines four complementary objectives: latent token prediction, cross-modal token alignment, SIGReg-based Gaussian regularization in the fused latent space, and a differentiable SWAP-test-based Fidelity Quantum Similarity (FQS) loss. Unlike pixel reconstruction methods, HQ-JEPA learns semantic representations directly in latent space and uses quantum state-overlap-based similarity as an additional regularization signal. We evaluate the pretrained encoder on GeoBench classification and segmentation tasks under linear probing and fine-tuning settings. Results show that HQ-JEPA achieves competitive and often superior performance over strong self-supervised and remote sensing foundation-model baselines, demonstrating the benefit of integrating predictive self-supervision, cross-modal geometric regularization, and quantum fidelity-based representation learning for remote sensing applications.

遥感多模态量子机器学习自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。