arXiv:2409.01314cs.CVcs.LG2024-09被引 1

通过解耦像素簇相似度,定位图像生成模型的薄弱区域。

Disentangling Mean Embeddings for Better Diagnostics of Image Generators

  • 用中心核对齐分解整体相似度为各像素簇贡献
  • 发现不同图像区域生成质量差异显著
  • 适合需要诊断生成缺陷的开发者与研究者

图像生成器的评估仍面临挑战,传统指标难以提供对图像特定区域的细致洞察。并非所有图像区域都能以相同难度被模型学习。本文提出一种新方法,通过中心核对齐将均值嵌入的余弦相似度解耦为各个像素簇余弦相似度的乘积,从而量化每个像素簇对整体生成性能的贡献。实验表明,该方法显著提升了生成结果的可解释性,并提高了在多种真实场景中识别模型生成异常区域的可能性。

原文摘要 · Abstract (English)

The evaluation of image generators remains a challenge due to the limitations of traditional metrics in providing nuanced insights into specific image regions. This is a critical problem as not all regions of an image may be learned with similar ease. In this work, we propose a novel approach to disentangle the cosine similarity of mean embeddings into the product of cosine similarities for individual pixel clusters via central kernel alignment. Consequently, we can quantify the contribution of the cluster-wise performance to the overall image generation performance. We demonstrate how this enhances the explainability and the likelihood of identifying pixel regions of model misbehavior across various real-world use cases.

图像生成可解释性模型诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。