用双先验提升无监督跨域图像检索效果
Text-Phase Synergy Network with Dual Priors for Unsupervised Cross-Domain Image Retrieval
- 引入文本与相位双重先验,增强语义指导
- 在多个基准上超越当前最优方法
- 适合做无监督跨域检索的研究者参考
本文研究无监督跨域图像检索(UCDIR),旨在不依赖标注数据的情况下,跨不同领域检索同类别图像。现有方法通常使用聚类生成的伪标签作为监督信号,用于域内表示学习和跨域特征对齐,但离散伪标签难以提供准确全面的语义引导,且对齐过程常忽略域特异性与语义信息的纠缠,导致表征语义退化,影响检索性能。为此,本文提出文本-相位协同网络(TPSNet)。首先,利用CLIP为每个域生成一组特定类别的提示词,称为域提示,作为提供更精确语义监督的文本先验;同时引入相位先验,即域不变的相位特征,嵌入原始图像表示中,以弥合域分布差异并保持语义完整性。通过双先验的协同作用,TPSNet在多个UCDIR基准上显著优于现有方法。
原文摘要 · Abstract (English)
This paper studies unsupervised cross-domain image retrieval (UCDIR), which aims to retrieve images of the same category across different domains without relying on labeled data. Existing methods typically utilize pseudo-labels, derived from clustering algorithms, as supervisory signals for intra-domain representation learning and cross-domain feature alignment. However, these discrete pseudo-labels often fail to provide accurate and comprehensive semantic guidance. Moreover, the alignment process frequently overlooks the entanglement between domain-specific and semantic information, leading to semantic degradation in the learned representations and ultimately impairing retrieval performance. This paper addresses the limitations by proposing a Text-Phase Synergy Network with Dual Priors(TPSNet). Specifically, we first employ CLIP to generate a set of class-specific prompts per domain, termed as domain prompt, serving as a text prior that offers more precise semantic supervision. In parallel, we further introduce a phase prior, represented by domain-invariant phase features, which is integrated into the original image representations to bridge the domain distribution gaps while preserving semantic integrity. Leveraging the synergy of these dual priors, TPSNet significantly outperforms state-of-the-art methods on UCDIR benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。