提出信息论框架优化扩散模型的引导策略,平衡生成一致性与多样性。
Information-Theoretic Classifier-Free Guidance with Adaptive Schedule Optimization

- 基于信息论构建可调节的引导调度机制
- 在ImageNet和COCO上实现更优的生成质量与多样性平衡
- 适合需要精细控制生成风格的研究者
扩散模型在图像、文生图及视频生成中表现优异,其条件生成常通过无分类器引导(CFG)控制。CFG通过提升引导权重增强条件一致性,但过强引导会降低多样性与分布覆盖范围。目前尚不清楚如何在反向轨迹中合理调控这一权衡,因为CFG诱导的分布并非由引导得分场决定的固定时间倾斜分布。为此,我们提出一种基于信息论的CFG调度优化框架。该方法利用干净终点参考来定义期望的一致性-覆盖权衡,并优化实际由引导采样器生成的分布以逼近此参考。我们推导出从样本和得分评估中估计目标的轨迹级公式,避免了显式密度估计。在ImageNet-512(EDM-XXL)和COCO(SD-XL)上的实验表明,学习到的调度策略在一致性和多样性权衡上优于恒定引导,并能选择性地在不同噪声水平分配引导强度。
原文摘要 · Abstract (English)
Diffusion models have achieved strong performance in image, text-to-image, and video generation, where conditional generation is often controlled by classifier-free guidance (CFG). CFG improves condition consistency by increasing a guidance weight, but stronger guidance typically reduces diversity and distributional coverage. It remains unclear how this consistency-coverage trade-off should be controlled across the reverse trajectory, since the distribution induced by CFG is not simply the fixed-time tilted distribution given by the guided score field. To address this issue, we propose an information-theoretic framework for CFG schedule optimization. Our approach uses a clean endpoint reference to specify the desired consistency-coverage trade-off, while optimizing the actual distribution induced by the guided sampler toward this reference. We derive trajectory-level formulas to estimate the objective from samples and score evaluations, avoiding explicit density estimation. On ImageNet-512 with EDM-XXL and COCO with SD-XL, the learned schedules achieve competitive or improved trade-offs over constant guidance and allocate guidance selectively across noise levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。