arXiv:2506.17237cs.CV2025-06

揭示扩散模型生成图像的底层计算机制,发现真实人脸处理更复杂。

Mechanistic Interpretability of Diffusion Models: Circuit-Level Analysis and Causal Validation

  • 通过干预实验分析扩散模型电路级结构,定位关键计算路径。
  • 真实人脸处理比合成图像复杂度高8.4%,注意力模式差异显著。
  • 识别出8类专注功能的注意力机制,可定向调控生成效果。

我们对扩散模型进行了定量电路级分析,揭示了图像生成过程中的计算路径与机制原理。通过对2000张合成图像和2000张CelebA人脸图像进行系统性干预实验,发现扩散架构在处理合成数据与自然图像时存在根本性算法差异。真实人脸处理所需电路的计算复杂度更高(复杂度比 = 1.084 ± 0.008,p < 0.001),其注意力机制表现出显著不同的专业化模式,各去噪时间步的熵偏移范围为0.015至0.166。我们识别出八类功能各异的注意力机制,分别承担边缘检测(熵 = 3.18 ± 0.12)、纹理分析(熵 = 4.16 ± 0.08)和语义理解(熵 = 2.67 ± 0.15)等任务。干预分析表明,在关键节点进行消融操作会导致性能下降25.6%至128.3%,为所识别电路功能提供了因果证据。这些发现为通过机制干预实现生成模型行为的量化理解与控制奠定了基础。

原文摘要 · Abstract (English)

We present a quantitative circuit-level analysis of diffusion models, establishing computational pathways and mechanistic principles underlying image generation processes. Through systematic intervention experiments across 2,000 synthetic and 2,000 CelebA facial images, we discover fundamental algorithmic differences in how diffusion architectures process synthetic versus naturalistic data distributions. Our investigation reveals that real-world face processing requires circuits with measurably higher computational complexity (complexity ratio = 1.084 plus/minus 0.008, p < 0.001), exhibiting distinct attention specialization patterns with entropy divergence ranging from 0.015 to 0.166 across denoising timesteps. We identify eight functionally distinct attention mechanisms showing specialized computational roles: edge detection (entropy = 3.18 plus/minus 0.12), texture analysis (entropy = 4.16 plus/minus 0.08), and semantic understanding (entropy = 2.67 plus/minus 0.15). Intervention analysis demonstrates critical computational bottlenecks where targeted ablations produce 25.6% to 128.3% performance degradation, providing causal evidence for identified circuit functions. These findings establish quantitative foundations for algorithmic understanding and control of generative model behavior through mechanistic intervention strategies.

扩散模型可解释性注意力机制因果分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。