arXiv:2607.20540cs.LGcs.AI2026-07被引 2

提出基于熵的噪声分配策略,让扩散模型更高效训练。

From Atoms to Entropy: Optimal Noise Allocation for Diffusion Training in the Convex Regime

论文配图:From Atoms to Entropy: Optimal Noise Allocation for Diffusion Training in the Convex Regime
图 1 · 摘自论文原文
  • 从信息论出发,推导出最优噪声水平分配的数学框架。
  • 发现最佳训练方案集中在有限个噪声级别上,且与熵增长速率相关。
  • 在图像和离散数据上验证,比传统方法更省时,适合大规模训练。

如何决定扩散模型在哪些噪声水平上训练以及训练强度?尽管这一选择至关重要,但现有噪声调度主要依赖启发式或经验调优。本文建立了一个通用统计框架,研究扩散训练中渐近最优的噪声水平分配。第一个主要结果针对全耦合情形,即不同时间点间信息可传递,在凸性或Polyak-Lojasiewicz型假设下,证明最优训练调度具有原子解,集中于有限个噪声水平。第二个结果将该框架特化到理想化的独立学习者情形,旨在建模神经网络中的时间专属性。在额外特征-噪声解耦条件下,通过随机矩阵分析得到一个信息论代理:解耦采样密度正比于生成熵率的平方根,即前向过程中条件熵的增长速率。我们在可直接优化耦合目标的受控环境中测试这些预测,包括狄拉克混合、低维流形和MNIST。结果表明,优化后的调度始终为有限支持,而平滑熵代理在神经网络模型中与原子最优高度一致,仅在完全耦合参数情形下失效,符合理论预期。随后在更大规模实验中评估熵调度,由于完整调度优化目前不可行,结果表明平方根熵调度可在离散域显著提升训练效率,并在连续图像上保持与标准EDM启发式相当的竞争力。

原文摘要 · Abstract (English)

How should a diffusion model decide which noise levels to train on, and how much? Despite the importance of this choice, current noise schedules are based largely on heuristics or empirical tuning. Here, we develop a general statistical framework for studying asymptotically optimal noise-level allocation in diffusion training. Our first main result concerns the fully coupled regime, where information can spread between different time points. Under convexity or Polyak-Lojasiewicz-type assumptions, we show that the optimized training schedule admits an atomic minimizer, concentrated on finitely many noise levels. Our second main result specializes this framework to an idealized independent-learner regime, intended to model temporal specialization in neural networks. Under an additional feature-noise decoupling condition, a random-matrix analysis leads to an information-theoretic proxy: the decoupled sampling density is proportional to the square root of the generative entropy rate, the rate at which conditional entropy grows along the forward process. We test these predictions in controlled settings where the coupled objective can be optimized directly, including Dirac mixtures, low-dimensional manifolds, and MNIST. In these settings, the optimized schedules are consistently finite-support, while the smooth entropic proxy closely tracks the atomic optimum in neural-network models and breaks down mainly in the fully coupled parametric case, as the theory suggests. We then evaluate the entropic schedule in larger-scale experiments, where full schedule optimization is currently intractable. The results indicate that square-root entropy scheduling can substantially improve training efficiency on discrete domains and remains competitive with standard EDM-style heuristics on continuous images.

扩散模型噪声调度信息论训练效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。