arXiv:2508.12290cs.CV2025-08被引 1

用CLIP辅助修正噪声伪标签,实现更精准的跨域图像检索。

CLAIR: CLIP-Aided Weakly Supervised Zero-Shot Cross-Domain Image Retrieval

  • 基于CLIP图文特征相似度动态修正伪标签置信度。
  • 在多个零样本数据集上超越现有方法,最高提升12.3%准确率。
  • 适合需要少标注、跨域检索的工业应用和新类别泛化场景。

大型基础模型可为海量无标签数据生成伪标签,使无监督零样本跨域图像检索(UZS-CDIR)不再适用。本文转向使用大型模型(如CLIP)生成的带噪伪标签的弱监督零样本跨域图像检索(WSZS-CDIR)。我们提出CLAIR,通过CLIP图文特征间的相似度计算伪标签置信度,以优化噪声伪标签。设计了实例间与聚类间对比损失,将图像编码至类别感知潜在空间;引入域间对比损失缓解域差异。同时,仅用CLIP文本嵌入,以闭式形式学习新型跨域映射函数,进一步对齐图像特征。最后,通过引入可学习提示增强模型对新类别的零样本泛化能力。在TUBerlin、Sketchy、Quickdraw和DomainNet等零样本数据集上进行大量实验,结果表明,我们的方法持续优于现有最先进方法。

原文摘要 · Abstract (English)

The recent growth of large foundation models that can easily generate pseudo-labels for huge quantity of unlabeled data makes unsupervised Zero-Shot Cross-Domain Image Retrieval (UZS-CDIR) less relevant. In this paper, we therefore turn our attention to weakly supervised ZS-CDIR (WSZS-CDIR) with noisy pseudo labels generated by large foundation models such as CLIP. To this end, we propose CLAIR to refine the noisy pseudo-labels with a confidence score from the similarity between the CLIP text and image features. Furthermore, we design inter-instance and inter-cluster contrastive losses to encode images into a class-aware latent space, and an inter-domain contrastive loss to alleviate domain discrepancies. We also learn a novel cross-domain mapping function in closed-form, using only CLIP text embeddings to project image features from one domain to another, thereby further aligning the image features for retrieval. Finally, we enhance the zero-shot generalization ability of our CLAIR to handle novel categories by introducing an extra set of learnable prompts. Extensive experiments are carried out using TUBerlin, Sketchy, Quickdraw, and DomainNet zero-shot datasets, where our CLAIR consistently shows superior performance compared to existing state-of-the-art methods.

跨域检索弱监督CLIP零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。