arXiv:2606.31517cs.IRcs.CV2026-06被引 1

用少量图文对实现高效跨模态检索,保留语义结构。

Unsupervised Data-Efficient Cross-Modal Retrieval with Global-Neighborhood Alignment Hashing

论文配图:Unsupervised Data-Efficient Cross-Modal Retrieval with Global-Neighborhood Alignment Hashing
图 1 · 摘自论文原文
  • 通过原型锚定全局对齐,将视觉语言模型的语义结构映射到二进制空间。
  • 在仅用少量数据下,检索准确率超越现有无监督方法。
  • 适合数据稀缺场景,如小样本跨模态应用。

与有监督跨模态哈希相比,无监督跨模态哈希(CMH)通过从无标注图文对中学习二进制编码,降低了对人工标注的依赖。然而,现有方法通常需要大规模图文对,收集成本高。为此,本文提出全局-邻域对齐哈希(GNAH),仅需少量图文对即可在紧凑的二进制海明空间中保留视觉-语言基础模型的语义结构。GNAH通过原型锚定全局对齐模块,从连续潜在空间捕获全局结构并传递至二进制空间;同时,通过对比随机邻域对齐模块,建模随机邻域关系,缓解对稀疏成对相关性的过拟合。大量实验表明,GNAH在数据受限条件下持续优于现有无监督跨模态检索方法,为实际应用场景提供了可行方案。

原文摘要 · Abstract (English)

Compared to supervised cross-modal hashing (CMH), unsupervised CMH reduces the reliance on manual labeling by learning binary codes from unlabeled image-text pairs. However, existing unsupervised CMH methods often rely on large-scale image-text pairs, which are costly to collect. To address this limitation, we propose Global-Neighborhood Alignment Hashing (GNAH), a novel approach that preserves the semantic structure of vision-language foundation models within a compact binary Hamming space using only a limited number of image-text pairs. Specifically, GNAH captures global structural information from the continuous latent space and transfers it into the binary Hamming space through a Prototype-Anchored Global Alignment module. In addition, GNAH extends conventional pairwise contrastive learning by modeling stochastic neighborhood relationships via a Contrastive Stochastic Neighborhood Alignment module, thereby alleviating overfitting to sparse pairwise correlations. Extensive experiments demonstrate that GNAH consistently outperforms existing unsupervised cross-modal retrieval methods under data-constrained settings, offering a practical solution for real-world CMH applications.

跨模态检索无监督学习哈希算法小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。