通过分析深层隐变量波动,提出无需训练的去噪框架,减少生成幻觉。
Unified Backbone Refinement for Diffusion Models via Internal-Latent Analysis

- 基于共享EMA准则检测深层隐变量异常波动
- 对异常区域实施针对性抑制,提升生成质量
- 适用于U-Net与Transformer模型,无需额外训练
扩散模型在多个领域表现卓越,其性能与参数化得分函数的去噪主干密切相关。本文提出一种阶段感知的系统性分析,发现深层低噪声隐变量中早期剧烈波动与伪影强相关。基于此,我们提出DUNE(Diffusion Unified Network refiNEr)——一种无需训练的优化框架,利用共享EMA准则检测深层低噪声内部隐变量中的突变,对检测出的条目实施主干特定抑制。尽管源自U-Net,该检测-抑制原则可自然扩展至基于Transformer的扩散模型,作用于深度自注意力块的隐变量。大量实验表明,DUNE在多种主干结构上均提升了保真度并减少了幻觉现象,为何时何地控制扩散主干提供了新视角。
原文摘要 · Abstract (English)
Diffusion models have achieved remarkable success across diverse domains, with performance closely related to the denoising backbones that parameterize the score function. In this paper, we present a systematic, phase-aware analysis of diffusion components and show that abrupt, early-stage fluctuations in deep latents are strongly associated with artifacts. Guided by these findings, we introduce DUNE (Diffusion Unified Network refiNEr), a training-free refinement framework that detects abrupt deviations in deep low-noise internal latents using a shared EMA-based criterion, and applies backbone-specific suppression to the detector-selected entries. Although derived from U-Net, the same detect-suppress principle extends naturally to Transformer-based diffusion models by acting on the latents of deep self-attention blocks. Extensive experiments across multiple backbones indicate that DUNE improves fidelity while reducing hallucinations, offering new insight into where and when diffusion backbones should be controlled.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。