arXiv:2412.14580cs.CV2024-12ICCV被引 17

用预训练扩散模型评估图像相似性,更贴近人眼判断。

DiffSim: Taming Diffusion Models for Evaluating Visual Similarity

  • 利用去噪U-Net的注意力层特征对齐,同时衡量外观与风格相似性。
  • 在风格和实例级别基准上均超越现有方法,接近人类视觉偏好。
  • 提出Sref和IP新基准,适合评估定制生成任务的视觉一致性。

扩散模型彻底改变了生成模型领域,使得评估定制化生成结果与参考输入之间的相似性变得至关重要。然而,传统感知相似性度量主要在像素和补丁层面操作,仅比较低层次的颜色与纹理,难以捕捉图像布局、物体姿态及语义内容等中层差异。基于对比学习的CLIP和自监督学习的DINO虽用于度量语义相似性,但过度压缩图像特征,无法充分评估外观细节。本文首次发现可利用预训练扩散模型来度量视觉相似性,并提出DiffSim方法,解决传统度量在定制生成任务中感知一致性捕捉不足的问题。通过对去噪U-Net注意力层特征进行对齐,DiffSim同时评估外观与风格相似性,表现出与人类视觉偏好更高的契合度。此外,我们引入Sref和IP两个新基准,分别用于评估风格和实例级别的视觉相似性。在多个基准上的综合评估表明,DiffSim达到当前最优性能,为生成模型的视觉连贯性测量提供可靠工具。

原文摘要 · Abstract (English)

Diffusion models have fundamentally transformed the field of generative models, making the assessment of similarity between customized model outputs and reference inputs critically important. However, traditional perceptual similarity metrics operate primarily at the pixel and patch levels, comparing low-level colors and textures but failing to capture mid-level similarities and differences in image layout, object pose, and semantic content. Contrastive learning-based CLIP and self-supervised learning-based DINO are often used to measure semantic similarity, but they highly compress image features, inadequately assessing appearance details. This paper is the first to discover that pretrained diffusion models can be utilized for measuring visual similarity and introduces the DiffSim method, addressing the limitations of traditional metrics in capturing perceptual consistency in custom generation tasks. By aligning features in the attention layers of the denoising U-Net, DiffSim evaluates both appearance and style similarity, showing superior alignment with human visual preferences. Additionally, we introduce the Sref and IP benchmarks to evaluate visual similarity at the level of style and instance, respectively. Comprehensive evaluations across multiple benchmarks demonstrate that DiffSim achieves state-of-the-art performance, providing a robust tool for measuring visual coherence in generative models.

扩散模型图像相似性视觉评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。