arXiv:2503.22782cs.CVcs.AI2025-03

用原型网络让扩散模型生成过程可解释,看清图像特征何时何地出现

Patronus: Interpretable Diffusion Models with Prototypes

  • 引入原型网络编码视觉块语义,追踪生成中模式演变
  • 在5个数据集上实现强生成性能与精准可解释性
  • 适合想理解或控制扩散模型生成逻辑的研究者

揭示基于扩散的生成模型的黑箱问题迫在眉睫,随着其应用持续扩展,其内部机制仍不透明。针对‘如何使扩散生成过程可解释’这一关键问题,我们提出Patronus——一种可解释的扩散模型,通过原型网络对视觉块进行语义编码,揭示了视觉模式在去噪过程中何时、何地涌现。该可解释性使我们能够检测由非期望相关性引发的捷径学习,并追踪语义在时间步上的演化路径。我们在四个自然图像数据集和一个医学影像数据集上评估Patronus,验证了其忠实的可解释性与强大的生成能力。本工作为通过原型基方法理解与操控扩散模型开辟了新路径。

原文摘要 · Abstract (English)

Uncovering the opacity of diffusion-based generative models is urgently needed, as their applications continue to expand while their underlying procedures largely remain a black box. With a critical question -- how can the diffusion generation process be interpreted and understood? -- we proposed Patronus, an interpretable diffusion model that incorporates a prototypical network to encode semantics in visual patches, revealing what visual patterns are modeled and where and when they emerge throughout denoising. This interpretability of Patronus provides deeper insights into the generative mechanism, enabling the detection of shortcut learning via unwanted correlations and the tracing of semantic emergence across timesteps. We evaluate Patronus on four natural image datasets and one medical imaging dataset, demonstrating both faithful interpretability and strong generative performance. With this work, we open new avenues for understanding and steering diffusion models through prototype-based interpretability.\\ Our code is available at https://github.com/nina-weng/patronus}{https://github.com/nina-weng/patronus.

可解释性扩散模型原型网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。