SUSS用可解释的概率模型评估图像相似性,更贴近人眼判断。
Structured Uncertainty Similarity Score (SUSS): Learning a Probabilistic, Interpretable, Perceptual Metric Between Images
- 将图像建模为带结构的多元正态分布,通过自监督学习捕捉感知特性
- 在多种失真下与人类判断高度一致,且能提供局部解释
- 适合需要可解释性的图像质量评估和生成任务
与人类视觉对齐的感知相似性评分对计算机视觉模型的训练与评估至关重要。深度感知损失(如LPIPS)虽效果良好,但依赖复杂非线性判别特征且不变性未知;手工设计指标(如SSIM)可解释性强,却遗漏关键感知属性。本文提出结构化不确定性相似性评分(SUSS):将每张图像建模为一组感知成分,每个成分由结构化的多元正态分布表示。这些成分通过生成式自监督方式训练,对人类不可察觉的增强具有高似然性。最终得分是各成分对数似然的加权和,权重来自人类感知数据集学习得到。与基于特征的方法不同,SUSS在像素空间中学习图像特定的残差线性变换,可通过去相关残差和采样进行透明分析。SUSS与人类感知判断高度一致,在多种失真类型下表现出强感知校准能力,并提供可定位的可解释性说明。我们进一步证明其在下游成像任务中作为感知损失时具有稳定优化行为和竞争力的表现。
原文摘要 · Abstract (English)
Perceptual similarity scores that align with human vision are critical for both training and evaluating computer vision models. Deep perceptual losses, such as LPIPS, achieve good alignment but rely on complex, highly non-linear discriminative features with unknown invariances, while hand-crafted measures like SSIM are interpretable but miss key perceptual properties. We introduce the Structured Uncertainty Similarity Score (SUSS); it models each image through a set of perceptual components, each represented by a structured multivariate Normal distribution. These are trained in a generative, self-supervised manner to assign high likelihood to human-imperceptible augmentations. The final score is a weighted sum of component log-probabilities with weights learned from human perceptual datasets. Unlike feature-based methods, SUSS learns image-specific linear transformations of residuals in pixel space, enabling transparent inspection through decorrelated residuals and sampling. SUSS aligns closely with human perceptual judgments, shows strong perceptual calibration across diverse distortion types, and provides localized, interpretable explanations of its similarity assessments. We further demonstrate stable optimization behavior and competitive performance when using SUSS as a perceptual loss for downstream imaging tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。