动态语义通道损失提升哈希检索精度,尤其在跨模态场景下表现更优。
DSCH-Loss: A Dynamic Semantic Channel Objective for Deep Semantic Hashing

- 设计动态大小与位置的语义通道,避免传统方法的优化断点问题。
- 在40项任务中35次超越现有方法,最长提升达1.75个百分点。
- 采用考虑重复得分的mAP评估,更准确反映实际检索效果。
为在高维数据空间中实现高效近似最近邻搜索,语义哈希生成短二进制哈希码的方法近年来受到广泛关注。基于深度学习的方法相比依赖人工特征工程的传统方法,具备更强的语义捕捉能力,并可在多种数据模态间实现数据驱动的语义哈希,生成共享汉明空间内的高质量跨模态哈希码。以往研究分析了该汉明空间特性,提出基于预定义语义通道的损失函数,其通道宽度固定且汉明距离由标签相似性导出。然而,该设定引入了损失函数景观中的不连续性,增加了优化难度。为此,本文提出新的动态语义通道哈希(DSCH)损失函数,通过动态调整语义通道的尺寸与位置,避免损失景观的不连续性。同时,推荐使用考虑重复情况的平均精度(tie-aware mAP)作为评估指标,以解决哈希码距离离散导致的样本检索排序模糊问题。在两个主流数据集上,采用两种不同模型架构的多组实验表明,使用DSCH目标训练的模型在40项任务中的35项上显著优于其他先进损失函数。在所有四个测试哈希码长度下,其性能均保持领先,相比第二佳方案最高提升1.75个百分点,结果稳定且具说服力。
原文摘要 · Abstract (English)
Semantic hashing methods for generating short binary hash codes that allow efficient approximate nearest neighbor search in high-dimensional data spaces have gained extensive consideration in recent years. Deep learning-based methods offer better semantic capturing capabilities than traditional approaches relying on manual feature engineering. Moreover, they enable a data-driven approach to semantic hashing across diverse data modalities, yielding high-quality cross-modal hash codes within a shared Hamming space. Previous work investigated the properties of this Hamming space and introduced a loss function based on predefined so-called semantic channels with fixed width and Hamming distances derived from label similarities. However, this formulation also introduced discontinuities into the loss landscape, complicating optimization. Based on these observations, we propose a newly designed loss function, Dynamic Semantic Channel Hashing (DSCH), using dynamically sized and positioned semantic channels in order to avoid loss landscape discontinuities. Furthermore, we endorse the use of tie-aware Mean Average Precision (mAP) as evaluation metric as it addresses the ambiguity in sample retrieval ordering, which emerges from the discreteness of hash code distances. Finally, multiple experimental settings conducted on two popular datasets and incorporating two different model architectures provide strong evidence that training using the DSCH objective outperforms training using other state-of-the-art loss functions. In a total of 35 out of 40 cross-modal and intra-modal retrieval tasks, models trained with DSCH achieve significantly higher tie-aware mAP scores across all four tested hash code lengths, showing compelling results across model architecture and used dataset. The mAP score uplifts are consistent and amount up to 1.75 percentage points compared to the respective second best.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。