arXiv:2510.20903cs.ITcs.LG2025-10NeurIPS被引 4

提出新方法提升扩散模型训练效率与理论可解释性。

Information Theoretic Learning for Diffusion Models with Warm Start

  • 基于信息论推导更紧的似然上界,放宽高斯噪声假设。
  • 在CIFAR-10和ImageNet上实现媲美或超越当前最优的负对数似然。
  • 适用于连续与离散数据,无需数据增强,适合注重理论严谨性的研究者。

以最大化模型似然为核心的生成模型在实际应用中日益流行。其中,基于扰动的方法支撑了众多强大的似然估计模型,但常面临收敛缓慢和理论理解不足的问题。本文推导出噪声驱动模型的更紧似然上界,从而提升最大似然学习的准确性和效率。核心洞见将经典的KL散度与费雪信息关系拓展至任意噪声扰动,突破高斯假设限制,支持结构化噪声分布。该框架可灵活使用随机噪声,自然建模传感器伪影、量化效应及数据分布平滑,同时兼容标准扩散训练流程。将扩散过程视为高斯信道,进一步表达数据与模型间的失配熵,证明所提目标上界为负对数似然(NLL)。实验表明,模型在CIFAR-10上取得竞争性NLL,ImageNet多分辨率下达当前最优结果,且无需数据增强;该框架还可自然推广至离散数据。

原文摘要 · Abstract (English)

Generative models that maximize model likelihood have gained traction in many practical settings. Among them, perturbation based approaches underpin many strong likelihood estimation models, yet they often face slow convergence and limited theoretical understanding. In this paper, we derive a tighter likelihood bound for noise driven models to improve both the accuracy and efficiency of maximum likelihood learning. Our key insight extends the classical KL divergence Fisher information relationship to arbitrary noise perturbations, going beyond the Gaussian assumption and enabling structured noise distributions. This formulation allows flexible use of randomized noise distributions that naturally account for sensor artifacts, quantization effects, and data distribution smoothing, while remaining compatible with standard diffusion training. Treating the diffusion process as a Gaussian channel, we further express the mismatched entropy between data and model, showing that the proposed objective upper bounds the negative log-likelihood (NLL). In experiments, our models achieve competitive NLL on CIFAR-10 and SOTA results on ImageNet across multiple resolutions, all without data augmentation, and the framework extends naturally to discrete data.

扩散模型信息论生成模型理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。