通过语义协同知识蒸馏,提升跨模态哈希的语义一致性。
Semantic-Cohesive Knowledge Distillation for Deep Cross-modal Hashing
- 将多标签重构为文本提示,作为新模态参与学习。
- 教师网络跨模态提取语义特征,生成可指导学生的哈希空间。
- 在两个基准数据集上优于现有方法,适合多模态检索场景。
近年来,深度监督跨模态哈希方法通过自监督方式学习语义信息取得了显著进展。然而,其仍存在关键缺陷:多标签语义提取过程未能显式与原始多模态数据交互,导致学习到的表示级语义与异构多模态数据不兼容,阻碍了模态间隙的弥合。为此,本文提出一种新型语义协同知识蒸馏方案SODA。具体地,将多标签信息引入为新的文本模态,并重构为一组真实标签提示,描述图像中呈现的语义,类似于文本模态。随后,设计一个跨模态教师网络,有效提炼图像与标签模态间的跨模态语义特征,从而为图像模态学习到一个良好映射的汉明空间。从某种意义上讲,该汉明空间可视为一种先验知识,指导跨模态学生网络的学习,并全面保留图像与文本模态间的语义相似性。在两个基准数据集上的大量实验表明,本模型显著优于当前最优方法。
原文摘要 · Abstract (English)
Recently, deep supervised cross-modal hashing methods have achieve compelling success by learning semantic information in a self-supervised way. However, they still suffer from the key limitation that the multi-label semantic extraction process fail to explicitly interact with raw multimodal data, making the learned representation-level semantic information not compatible with the heterogeneous multimodal data and hindering the performance of bridging modality gap. To address this limitation, in this paper, we propose a novel semantic cohesive knowledge distillation scheme for deep cross-modal hashing, dubbed as SODA. Specifically, the multi-label information is introduced as a new textual modality and reformulated as a set of ground-truth label prompt, depicting the semantics presented in the image like the text modality. Then, a cross-modal teacher network is devised to effectively distill cross-modal semantic characteristics between image and label modalities and thus learn a well-mapped Hamming space for image modality. In a sense, such Hamming space can be regarded as a kind of prior knowledge to guide the learning of cross-modal student network and comprehensively preserve the semantic similarities between image and text modality. Extensive experiments on two benchmark datasets demonstrate the superiority of our model over the state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。