arXiv:2508.13628cs.CV2025-08被引 1

通过迭代误差优化,让扩散模型生成更稳定高质量样本

DiffIER: Optimizing Diffusion Models with Iterative Error Reduction

  • 在推理每一步迭代减少累积误差,提升生成质量
  • 指导权重敏感性显著降低,多任务表现优于基线
  • 无需修改模型,可插拔式增强各类生成任务

扩散模型在无分类器引导(CFG)下展现出强大的生成能力,但其输出质量对指导权重极为敏感。本文识别出训练-推理间的“间隙”是导致条件生成性能下降和权重敏感性的根源。通过量化推理阶段的累积误差,我们建立了指导权重选择与最小化该间隙之间的关联。为此,提出DiffIER——一种基于优化的迭代误差减小方法,在推理的每个步骤中优化误差,有效缓解该间隙。实验表明,该方法在文本到图像、图像超分辨率和文本到语音生成等任务中均显著优于基线,且具有良好的泛化能力,适用于多种生成场景。

原文摘要 · Abstract (English)

Diffusion models have demonstrated remarkable capabilities in generating high-quality samples and enhancing performance across diverse domains through Classifier-Free Guidance (CFG). However, the quality of generated samples is highly sensitive to the selection of the guidance weight. In this work, we identify a critical ``training-inference gap'' and we argue that it is the presence of this gap that undermines the performance of conditional generation and renders outputs highly sensitive to the guidance weight. We quantify this gap by measuring the accumulated error during the inference stage and establish a correlation between the selection of guidance weight and minimizing this gap. Furthermore, to mitigate this gap, we propose DiffIER, an optimization-based method for high-quality generation. We demonstrate that the accumulated error can be effectively reduced by an iterative error minimization at each step during inference. By introducing this novel plug-and-play optimization framework, we enable the optimization of errors at every single inference step and enhance generation quality. Empirical results demonstrate that our proposed method outperforms baseline approaches in conditional generation tasks. Furthermore, the method achieves consistent success in text-to-image generation, image super-resolution, and text-to-speech generation, underscoring its versatility and potential for broad applications in future research.

扩散模型生成优化条件生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。