让图像相似度评估可按文本提示灵活调整,更贴近人类感知。
The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric

- 用多维度文本提示控制相似性判断,支持形状、颜色等不同感知角度。
- 在图像三元组上构建大规模人类标注数据集,覆盖多种语义相似性。
- 新指标TPIPS更贴近人眼判断,适用于生成模型评估与精准检索。
人类对图像相似性的判断具有情境依赖性:例如两张图可能形状相似但色彩不同。现有感知相似性度量将这些差异简化为单一标量值,无法根据特定方面进行条件化。为填补这一空白,我们构建了一个大规模的人类相似性判断数据集,包含图像三元组,每个三元组在多个自由形式的语义层面(如形状、颜色、纹理等)进行了标注。对一系列前沿视觉-语言模型(VLMs)进行基准测试,发现其性能与人类共识之间存在显著差距。利用该数据集,我们微调一个VLM,提出文本提示式图像感知相似性度量(TPIPS),可根据指定的文本提示捕捉不同维度的视觉相似性。实验表明,TPIPS与人类感知高度一致,并在训练分布外具备良好泛化能力。此外,TPIPS可实现文本引导检索、组合式搜索及生成模型的细粒度评估。代码、数据与训练模型已公开于https://peterwang512.github.io/TPIPS。
原文摘要 · Abstract (English)
Human visual similarity judgments are context-dependent. For example, two images may be similar in shape but distinct in color. Existing perceptual similarity metrics, however, collapse these nuances into a single scalar value, offering no mechanism to condition on specific aspects. To bridge this gap, we introduce a large-scale dataset of human similarity judgments over image triplets, where each triplet is annotated across multiple, free-form semantic aspects of similarity. Benchmarking a broad range of frontier vision-language models (VLMs) reveals a considerable performance gap compared to human annotators' consensus. Leveraging our data, we fine-tune a VLM to produce our Text-Prompted Image Perceptual Similarity (TPIPS) metric, capturing multiple senses of visual similarity depending on the specified text prompt. We demonstrate that TPIPS aligns more closely with human perception and generalizes reliably beyond the training distribution. Finally, we show that TPIPS unlocks new capabilities in text-guided retrieval, compositional search, and the fine-grained evaluation of generative models. Our code, data, and trained models are at https://peterwang512.github.io/TPIPS
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。