arXiv:2609.08253cs.LG2026-09

通过谱表示对齐提升扩散模型训练与生成质量

Revisiting Spectral Representations in Generative Diffusion Models

论文配图:Revisiting Spectral Representations in Generative Diffusion Models
图 1 · 摘自论文原文
  • 利用随机扰动核构建自监督谱表示,实现隐状态对齐
  • 在图像和3D点云上均提升生成质量,效果稳定
  • 揭示了谱对齐与扩散得分蒸馏的数学等价性,适合相关研究者

扩散模型在多种生成任务中表现优异。近期研究表明,在扩散网络的隐状态上施加表示对齐可加速训练收敛并提升采样质量,但其内在机制尚不明确。本文从扰动核的统一视角,探讨自监督谱表示学习与扩散生成模型的关联:扩散过程通过高斯核反向注入噪声生成样本;谱表示则源于随机扰动核诱导的正负关系对比。受此启发,我们提出一种自监督谱表示对齐方法以促进扩散模型训练,并从几何角度阐明联合谱学习如何助力训练。进一步发现,谱对齐优化目标在表示空间中等价于扩散得分蒸馏。基于此,我们在扩散训练目标中引入谱正则项,显著提升了多个数据集上的生成性能。实验涵盖图像与3D点云,结果一致表明生成质量改善。代码已开源:https://github.com/yuehaowang/spectral-reg-diffusion。

原文摘要 · Abstract (English)

Diffusion models have shown remarkable performance on diverse generation tasks. Recent work finds that imposing representation alignment on the hidden states of diffusion networks can both facilitate training convergence and enhance sampling quality, yet the mechanism driving this synergy remains insufficiently understood. In this paper, we investigate the connection between self-supervised spectral representation learning and diffusion generative models through a shared perspective on perturbation kernels. On the diffusion side, samples (e.g., images, videos) are produced by reversing a stochastic noise-injection process specified by Gaussian kernels; on the spectral representation side, spectral embeddings emerge from contrasting positive and negative relations induced by random perturbation kernels. Motivated by this, we propose a self-supervised spectral representation alignment method to facilitate diffusion model training. In addition, we clarify how joint spectral learning can benefit diffusion training from a geometric perspective. Furthermore, we find that the optimization of the spectral alignment objective is in an equivalent form of diffusion score distillation in the representation space. Building on these findings, we integrate a spectral regularizer into diffusion training objectives to improve the performance of diffusion models on multiple datasets. Experiments across images and 3D point clouds show consistent gains in generation quality. Code is released at https://github.com/yuehaowang/spectral-reg-diffusion.

扩散模型谱表示生成模型正则化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。