arXiv:2410.07783cs.CV2024-10中稿 · 31st International…被引 3

用CLIP提升多模态哈希检索精度,效果显著。

CLIP Multi-modal Hashing for Multimedia Retrieval

  • 基于CLIP提取文本与视觉特征并融合生成哈希码
  • 在多个数据集上相比顶尖方法最高提升8.38% mAP
  • 适合需要高精度多模态检索的工程应用

多模态哈希方法广泛用于多媒体检索,可通过融合多源数据生成二进制哈希码。然而,传统方法中的独立骨干网络特征表达能力有限,且未在大规模无监督多模态数据上联合预训练,导致检索准确率较低。为此,本文提出一种新型的CLIP多模态哈希(CLIPMH)方法。该方法利用CLIP框架提取文本和视觉特征,并将其融合生成哈希码。由于各模态特征得到增强,所提方法在多模态哈希检索性能上实现显著提升。实验表明,相较于当前最先进的无监督与有监督多模态哈希方法,CLIPMH在平均精度均值(mAP)上可实现最大8.38%的提升。

原文摘要 · Abstract (English)

Multi-modal hashing methods are widely used in multimedia retrieval, which can fuse multi-source data to generate binary hash code. However, the individual backbone networks have limited feature expression capabilities and are not jointly pre-trained on large-scale unsupervised multi-modal data, resulting in low retrieval accuracy. To address this issue, we propose a novel CLIP Multi-modal Hashing (CLIPMH) method. Our method employs the CLIP framework to extract both text and vision features and then fuses them to generate hash code. Due to enhancement on each modal feature, our method has great improvement in the retrieval performance of multi-modal hashing methods. Compared with state-of-the-art unsupervised and supervised multi-modal hashing methods, experiments reveal that the proposed CLIPMH can significantly improve performance (a maximum increase of 8.38% in mAP).

多模态哈希CLIP检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。