arXiv:2603.02447cs.LG2026-03被引 1

通过频域正则化提升扩散模型生成质量,尤其改善高分辨率细节。

Spectral Regularization for Diffusion Models

  • 在训练中加入可微的傅里叶与小波域损失,不改动模型结构
  • 高分辨率图像和音频生成质量显著提升,细粒度结构更真实
  • 兼容主流扩散模型,计算开销极低,适合实际应用

扩散模型通常采用忽略自然信号频谱与多尺度结构的点对点重建目标进行训练。我们提出一种损失级频谱正则化框架,在不修改扩散过程、模型架构或采样步骤的前提下,引入可微的傅里叶域与小波域损失。该正则化作为软归纳偏置,促使生成样本具备合理的频率平衡与连贯的多尺度结构。方法兼容DDPM、DDIM与EDM框架,计算开销可忽略。在图像与音频生成任务上的实验表明,样本质量持续提升,尤其在高分辨率无条件数据集上,细粒度结构建模效果最为显著。

原文摘要 · Abstract (English)

Diffusion models are typically trained using pointwise reconstruction objectives that are agnostic to the spectral and multi-scale structure of natural signals. We propose a loss-level spectral regularization framework that augments standard diffusion training with differentiable Fourier- and wavelet-domain losses, without modifying the diffusion process, model architecture, or sampling procedure. The proposed regularizers act as soft inductive biases that encourage appropriate frequency balance and coherent multi-scale structure in generated samples. Our approach is compatible with DDPM, DDIM, and EDM formulations and introduces negligible computational overhead. Experiments on image and audio generation demonstrate consistent improvements in sample quality, with the largest gains observed on higher-resolution, unconditional datasets where fine-scale structure is most challenging to model.

扩散模型频谱正则图像生成音频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。