arXiv:2504.17804cs.CV2025-04被引 2

用频谱基函数线性组合生成图像,可解释性强且训练稳定。

Spectral Dictionary Learning for Generative Image Modeling

  • 将图像转为一维信号,用可调频率/相位/振幅的基函数线性重构。
  • 在CIFAR-10上重建质量与感知保真度达竞品水平,训练更稳定。
  • 适合需要可控图像合成与频域操作的场景,如纹理编辑或分析。

我们提出一种新型频谱生成模型用于图像合成,彻底区别于常见的变分、对抗和扩散范式。图像经展平为一维信号后,被重构为一组学习得到的频谱基函数的线性组合,每个基函数显式地由频率、相位和振幅参数化。模型联合学习一个具有时变调制的全局频谱字典及每张图像的混合系数,量化各频谱成分的贡献。随后,对这些混合系数拟合一个简单概率模型,通过从隐空间采样即可确定性生成新图像。该框架利用确定性字典学习,相比依赖随机推断或对抗训练的方法,具有更高可解释性和物理意义。此外,通过短时傅里叶变换(STFT)计算的频域损失函数,确保合成图像同时捕捉全局结构和细粒度频谱细节,如纹理与边缘信息。在CIFAR-10基准上的实验表明,该方法不仅在重建质量与感知保真度上表现优异,且具备更好的训练稳定性和计算效率。这种新型生成模型为可控合成开辟了新路径,因所学频谱字典可直接操控图像内在频率内容,提升可解释性,并有望拓展至图像编辑与分析等新应用。

原文摘要 · Abstract (English)

We propose a novel spectral generative model for image synthesis that departs radically from the common variational, adversarial, and diffusion paradigms. In our approach, images, after being flattened into one-dimensional signals, are reconstructed as linear combinations of a set of learned spectral basis functions, where each basis is explicitly parameterized in terms of frequency, phase, and amplitude. The model jointly learns a global spectral dictionary with time-varying modulations and per-image mixing coefficients that quantify the contributions of each spectral component. Subsequently, a simple probabilistic model is fitted to these mixing coefficients, enabling the deterministic generation of new images by sampling from the latent space. This framework leverages deterministic dictionary learning, offering a highly interpretable and physically meaningful representation compared to methods relying on stochastic inference or adversarial training. Moreover, the incorporation of frequency-domain loss functions, computed via the short-time Fourier transform (STFT), ensures that the synthesized images capture both global structure and fine-grained spectral details, such as texture and edge information. Experimental evaluations on the CIFAR-10 benchmark demonstrate that our approach not only achieves competitive performance in terms of reconstruction quality and perceptual fidelity but also offers improved training stability and computational efficiency. This new type of generative model opens up promising avenues for controlled synthesis, as the learned spectral dictionary affords a direct handle on the intrinsic frequency content of the images, thus providing enhanced interpretability and potential for novel applications in image manipulation and analysis.

图像生成频谱建模可解释性字典学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。