arXiv:2601.01045cs.LG2026-01

通过信息势能控制扩散生成模型的粗粒度特征,实现稳定可控的图像生成。

Coarse-Grained Kullback--Leibler Control of Diffusion-Based Generative AI

  • 基于信息势能设计投影反向扩散算法,显式调控图像块级统计量。
  • 数值实验显示块质量误差和容忍势能均在预设范围内,且保持像素级质量。
  • 适合需要结构化特征控制的生成任务,如图像编辑与风格迁移。

扩散模型与基于得分的生成模型为从噪声合成高质量图像提供了强大框架。然而,目前尚无理论描述粗粒度量(如图像分块后的块内强度或类别比例)在反向扩散动态中的保持与演化过程。本文将作者先前提出的用于非遍历马尔可夫过程的信息论李雅普诺夫函数 V 移植到生成模型的反向扩散过程,提出一种由泄漏容忍势能 V-delta 投影的反向扩散方案(称为 V-delta 投影反向扩散)。该方法在小泄漏条件下,使 V-delta 近似为李雅普诺夫函数。通过一个包含块常数图像与简化反向核的玩具模型,数值验证表明:所提方法能将块质量误差与泄漏容忍势能维持在预设容差内,同时达到与非投影动态相当的像素级精度与视觉质量。本研究将生成采样重新诠释为信息势能从噪声到数据的下降过程,并为具有粗粒度控制能力的反向扩散过程提供了设计原则。

原文摘要 · Abstract (English)

Diffusion models and score-based generative models provide a powerful framework for synthesizing high-quality images from noise. However, there is still no satisfactory theory that describes how coarse-grained quantities, such as blockwise intensity or class proportions after partitioning an image into spatial blocks, are preserved and evolve along the reverse diffusion dynamics. In previous work, the author introduced an information-theoretic Lyapunov function V for non-ergodic Markov processes on a state space partitioned into blocks, defined as the minimal Kullback-Leibler divergence to the set of stationary distributions reachable from a given initial condition, and showed that a leak-tolerant potential V-delta with a prescribed tolerance for block masses admits a closed-form expression as a scaling-and-clipping operation on block masses. In this paper, I transplant this framework to the reverse diffusion process in generative models and propose a reverse diffusion scheme that is projected by the potential V-delta (referred to as the V-delta projected reverse diffusion). I extend the monotonicity of V to time-inhomogeneous block-preserving Markov kernels and show that, under small leakage and the V-delta projection, V-delta acts as an approximate Lyapunov function. Furthermore, using a toy model consisting of block-constant images and a simplified reverse kernel, I numerically demonstrate that the proposed method keeps the block-mass error and the leak-tolerant potential within the prescribed tolerance, while achieving pixel-wise accuracy and visual quality comparable to the non-projected dynamics. This study reinterprets generative sampling as a decrease of an information potential from noise to data, and provides a design principle for reverse diffusion processes with explicit control of coarse-grained quantities.

扩散模型生成模型信息论可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。