arXiv:2606.15236cs.CV2026-06被引 1

通过频谱强制显式分离信号与噪声,提升像素空间扩散模型效率

Show the Signal, Hide the Noise: Spectral Forcing for Pixel-Space Diffusion

论文配图:Show the Signal, Hide the Noise: Spectral Forcing for Pixel-Space Diffusion
图 1 · 摘自论文原文
  • 引入无参频谱强制算子,随时间动态低通滤波输入图像
  • 在ImageNet-256上持续提升FID与Inception Score,训练全程有效
  • 适用于高频噪声为主的图像,尤其适合粗粒度分块生成任务

像素空间扩散模型在全带宽噪声图像上训练,但去噪器可利用的有效信号具有强频率依赖性。在修正流扩散与自然图像幂律谱下,各时刻t的带宽边界 $k^{*}(t) = (1-t)^{-2/α}$ 将低频信号区与高频噪声区分开。我们发现此隐含的从粗到精结构不仅描述现象,更引发容量分配问题:标准去噪器需自行识别移动边界,可能在最优预测退化为确定性基线的频时区域浪费计算。为此提出频谱强制,一种参数无关、时间条件的2D-DCT低通算子,作用于补丁嵌入前的噪声输入,其截止频率随扩散时间单调上升,至数据终点变为恒等变换。受控合成实验表明该算子在粗粒度分块与高频以噪声为主的图像中表现最佳。在ImageNet-256上,JiT-700M/32模型经频谱强制后,在不同训练阶段均稳定提升FID与Inception Score;在更细粒度分块下仍具竞争力。进一步将不变算子插入统一文生图模型SenseNova-U1,显著改善DPG-Bench与GenEval指标,证明输入侧频谱先验可迁移至跨类别生成。结果表明,通过显式展示信号、隐藏噪声,可实现容量高效的像素空间扩散。

原文摘要 · Abstract (English)

Pixel-space diffusion models are trained on full-bandwidth noisy images, yet the useful signal available to the denoiser is strongly frequency dependent. Under rectified-flow diffusion and natural-image power-law spectra, the per-band data-to-noise contour $k^{*}(t) = (1-t)^{-2/α}$ separates a signal-bearing low-frequency region from a noise-dominated high-frequency region at each time $t$. We show that this implicit coarse-to-fine structure is not merely descriptive: it induces a capacity-allocation problem. A standard pixel-space denoiser must discover the moving bandwidth boundary internally and can spend computation on frequency-time regions where the optimal prediction collapses to deterministic baselines rather than data-distribution modeling. To make this boundary explicit, we introduce Spectral Forcing, a parameter-free, time-conditional 2D-DCT low-pass operator applied to the noisy input before the patch embedder. Its cutoff expands monotonically with the diffusion time and becomes the identity at the data endpoint. Through controlled synthetic experiments, we identify the regime in which the operator is beneficial: coarse patch tokenization and data whose high-frequency content is predominantly noise rather than essential signal. On ImageNet-256 with JiT-700M/32, Spectral Forcing consistently improves both FID and Inception Score across different training epochs, demonstrating robust gains throughout training; at finer tokenization, the spectral forcing is still competitive. We further insert the unchanged operator into SenseNova-U1, a unified text-to-image model, where it improves DPG-Bench and GenEval, showing that the input-side spectral prior transfers beyond class-conditional generation. These results suggest a route to capacity-efficient pixel-space diffusion by showing the signal and hiding the noise.

扩散模型频谱分析生成效率像素空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。