arXiv:2605.28900cs.LG2026-05

通过谱特征实现无需重训练的高效扩散模型控制。

Spectral Guidance for Flexible and Efficient Control of Diffusion Models

论文配图:Spectral Guidance for Flexible and Efficient Control of Diffusion Models
图 1 · 摘自论文原文
  • 利用生成过程的内在几何结构提取关键控制特征。
  • 在CIFAR-10上比最强无训练基线提升37个百分点,采样速度提升4倍。
  • 支持标签、CLIP嵌入、掩码等多种控制方式,且可定位最佳引导时机。

我们提出Spectral Guidance框架,通过利用生成过程的内在几何结构来控制扩散模型。随着数据逐步被噪声污染,仅有少数特征对控制仍具信息量。我们将这些特征定义为条件期望算子的奇异函数,并通过自监督目标学习它们。一旦恢复,该基底可将任意引导信号(如标签、CLIP嵌入或掩码)直接投影到采样轨迹上。该方法无需重训练或采样时反向传播去噪器,即可实现稳定、高保真的控制。实验表明,在CIFAR-10上相比最强无训练基线,条件准确率提升37个百分点,采样速度加快4倍。此外,同一表示既支持标签与CLIP引导,也支持无需额外模型的空间控制(如掩码引导)。最后,该框架揭示了生成过程中的相变现象,精准定位有效引导的最佳时间窗口。

原文摘要 · Abstract (English)

We introduce Spectral Guidance, a framework for controlling diffusion models by leveraging the intrinsic geometry of the generative process. As data is progressively corrupted by noise, only a small number of features remain informative for control. We characterize them as the singular functions of a conditional expectation operator and show that they can be learned via a self-supervised objective. Once recovered, this basis enables the projection of arbitrary guidance signals, such as labels, CLIP embeddings, or masks, directly onto the sampling trajectory. This approach allows for stable, high-fidelity control without retraining or denoiser backpropagation during sampling. Empirically, we improve conditional accuracy on CIFAR-10 by 37 percentage points over the strongest training-free baseline while offering $4\times$ faster sampling. Moreover, the same representations that support label and CLIP guidance also enable spatial control, such as mask-based guidance, without auxiliary models. Finally, our framework reveals a phase transition in the generative process, pinpointing the optimal time window for effective guidance.

扩散模型可控生成谱方法无训练控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。