arXiv:2504.03821cs.CVcs.LG2025-04被引 8

用小波与傅里叶混合频域建模,提升图像生成的细节与结构控制。

A Hybrid Wavelet-Fourier Method for Next-Generation Conditional Diffusion Models

  • 结合小波与傅里叶变换,在混合频域中进行扩散过程
  • 在CIFAR-10、CelebA-HQ等数据集上FID优于基线模型
  • 适合需要高精度纹理和全局一致性的图像生成任务

我们提出一种新型生成建模框架——小波-傅里叶扩散模型(Wavelet-Fourier-Diffusion),将扩散范式适配到混合频域表示,以生成高质量、高保真图像并改善空间定位。与仅依赖像素空间加性噪声的传统扩散模型不同,本方法采用多变换策略,融合小波子带分解与部分傅里叶步骤,在前向和反向扩散过程中逐步降级并重建图像于混合谱域。通过引入小波的空间局部化能力补充传统傅里叶分析,模型能更有效捕捉全局结构与细粒度特征。我们进一步通过交叉注意力集成嵌入或条件特征,扩展至条件图像生成。在CIFAR-10、CelebA-HQ及条件ImageNet子集上的实验表明,该方法在弗雷歇初始距离(FID)和初始分数(IS)上达到或超过基线扩散模型与先进GANs的表现。同时验证了混合频域表示对全局连贯性与细粒度纹理合成的控制优势,为多尺度生成建模开辟新方向。

原文摘要 · Abstract (English)

We present a novel generative modeling framework,Wavelet-Fourier-Diffusion, which adapts the diffusion paradigm to hybrid frequency representations in order to synthesize high-quality, high-fidelity images with improved spatial localization. In contrast to conventional diffusion models that rely exclusively on additive noise in pixel space, our approach leverages a multi-transform that combines wavelet sub-band decomposition with partial Fourier steps. This strategy progressively degrades and then reconstructs images in a hybrid spectral domain during the forward and reverse diffusion processes. By supplementing traditional Fourier-based analysis with the spatial localization capabilities of wavelets, our model can capture both global structures and fine-grained features more effectively. We further extend the approach to conditional image generation by integrating embeddings or conditional features via cross-attention. Experimental evaluations on CIFAR-10, CelebA-HQ, and a conditional ImageNet subset illustrate that our method achieves competitive or superior performance relative to baseline diffusion models and state-of-the-art GANs, as measured by Fréchet Inception Distance (FID) and Inception Score (IS). We also show how the hybrid frequency-based representation improves control over global coherence and fine texture synthesis, paving the way for new directions in multi-scale generative modeling.

扩散模型频域建模图像生成小波变换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。