arXiv:2604.11539cs.CVcs.AI2026-04被引 1

让图像检索根据用户条件动态调整相似度判断

CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding Space

论文配图:CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding Space
图 1 · 摘自论文原文
  • 将视觉语言模型嵌入空间改造成文本条件化的相似度空间
  • 在多个数据集上实现更高检索准确率和更快计算速度
  • 适合需要多条件灵活检索的场景,如个性化推荐

人类对视觉相似性的感知具有自适应性和主观性,取决于用户的兴趣与关注点。然而,大多数图像检索系统依赖固定、单一的度量标准,无法同时融合多种条件。为此,我们提出CLAY,一种无需额外训练的自适应相似度计算方法,将预训练视觉-语言模型(VLMs)的嵌入空间重构为文本条件化的相似度空间。该设计分离了文本条件处理与视觉特征提取过程,支持使用固定视觉嵌入实现高效且多条件的检索。我们还构建了合成评估数据集CLAY-EVAL,用于在多样化的条件检索设置下进行全面评估。在标准数据集和所提数据集上的实验表明,相比以往方法,CLAY在检索准确率和计算效率方面均表现优异。

原文摘要 · Abstract (English)

Human perception of visual similarity is inherently adaptive and subjective, depending on the users' interests and focus. However, most image retrieval systems fail to reflect this flexibility, relying on a fixed, monolithic metric that cannot incorporate multiple conditions simultaneously. To address this, we propose CLAY, an adaptive similarity computation method that reframes the embedding space of pretrained Vision-Language Models (VLMs) as a text-conditional similarity space without additional training. This design separates the textual conditioning process and visual feature extraction, allowing highly efficient and multi-conditioned retrieval with fixed visual embeddings. We also construct a synthetic evaluation dataset CLAY-EVAL, for comprehensive assessment under diverse conditioned retrieval settings. Experiments on standard datasets and our proposed dataset show that CLAY achieves high retrieval accuracy and notable computational efficiency compared to previous works.

图像检索视觉语言模型条件相似度多条件查询

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。