arXiv:2504.07008cs.CV2025-04被引 3

发现扩散模型中间特征含位置编码和异常,提醒下游任务需谨慎使用

Latent Diffusion U-Net Representations Contain Positional Embeddings and Anomalies

  • 分析稳定扩散模型的表征相似性与范数,揭示内部机制
  • 发现中间层存在可学习的位置嵌入,且有高相似度角落伪影
  • 识别出高范数异常特征,提示表征鲁棒性存疑,适合研究表征可靠性者阅读

扩散模型在生成逼真图像方面表现出色,激发了将其表征用于下游任务的兴趣。为更好地理解这些表征的鲁棒性,我们采用表征相似性和范数分析方法,研究了主流的Stable Diffusion模型。研究发现三个现象:(1) 中间表征中存在可学习的位置嵌入;(2) 出现高相似度的角落伪影;(3) 存在异常的高范数伪影。这些发现表明,在将扩散模型表征用于对特征鲁棒性要求高的下游任务之前,必须进一步探究其内在属性。项目页面:https://jonasloos.github.io/sd-representation-anomalies

原文摘要 · Abstract (English)

Diffusion models have demonstrated remarkable capabilities in synthesizing realistic images, spurring interest in using their representations for various downstream tasks. To better understand the robustness of these representations, we analyze popular Stable Diffusion models using representational similarity and norms. Our findings reveal three phenomena: (1) the presence of a learned positional embedding in intermediate representations, (2) high-similarity corner artifacts, and (3) anomalous high-norm artifacts. These findings underscore the need to further investigate the properties of diffusion model representations before considering them for downstream tasks that require robust features. Project page: https://jonasloos.github.io/sd-representation-anomalies

扩散模型表征分析位置编码异常检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。