通过设计频谱各向异性噪声,让扩散模型更懂数据关键特征。
Learning What Matters: Steering Diffusion via Spectrally Anisotropic Forward Noise
- 用频率对角的结构化噪声替代传统均匀噪声,引导生成过程。
- 在多个视觉数据集上优于标准扩散模型,且能忽略特定频段的干扰。
- 适合需要精准控制生成偏差的研究者,如图像修复与抗噪生成。
扩散概率模型虽表现强劲,但其归纳偏置大多隐含。本文引入一种各向异性噪声算子,将频谱结构化的协方差替代原有各向同性前向协方差,构建显式归纳偏置。该方法统一了带通掩码与幂律加权,可选择性增强或抑制特定频段,同时保持前向过程为高斯分布。我们定义了频谱各向异性高斯扩散(SAGD),推导出其得分关系,并证明在全支持条件下,学习到的得分随 t→0 收敛至真实数据得分;同时,各向异性重塑了从噪声到数据的概率流路径。实验表明,所提方法在多个视觉数据集上优于标准扩散模型,且具备选择性忽略特定频段已知退化的能力。结果表明,精心设计的各向异性前向噪声为调节扩散模型归纳偏置提供了一种简洁而严谨的手段。
原文摘要 · Abstract (English)
Diffusion Probabilistic Models (DPMs) have achieved strong generative performance, yet their inductive biases remain largely implicit. In this work, we aim to build inductive biases into the training and sampling of diffusion models to better accommodate the target distribution of the data to model. We introduce an anisotropic noise operator that shapes these biases by replacing the isotropic forward covariance with a structured, frequency-diagonal covariance. This operator unifies band-pass masks and power-law weightings, allowing us to emphasize or suppress designated frequency bands, while keeping the forward process Gaussian. We refer to this as Spectrally Anisotropic Gaussian Diffusion (SAGD). In this work, we derive the score relation for anisotropic forward covariances and show that, under full support, the learned score converges to the true data score as $t\!\to\!0$, while anisotropy reshapes the probability-flow path from noise to data. Empirically, we show the induced anisotropy outperforms standard diffusion across several vision datasets, and enables selective omission: learning while ignoring known corruptions confined to specific bands. Together, these results demonstrate that carefully designed anisotropic forward noise provides a simple, yet principled, handle to tailor inductive bias in DPMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。