提出新方法提升文本生成图像的效率与质量
Towards a Golden Classifier-Free Guidance Path via Foresight Fixed Point Iterations
- 将条件生成重构为固定点迭代,找到最优路径
- 实验显示新方法在多个模型上图像质量更高、速度更快
- 适合关注扩散模型优化的研究者与开发者
Classifier-Free Guidance(CFG)是文本到图像扩散模型的核心组件,理解并改进其工作机制仍是研究重点。现有方法基于不同的理论解释,限制了设计空间并模糊了关键决策。为此,我们提出统一视角,将条件引导重新建模为固定点迭代,旨在寻找一个黄金路径:在条件与无条件生成下,潜在表示产生一致输出。我们证明,CFG及其变体属于单步短区间迭代的特例,理论上存在效率不足。为此,我们引入前瞻性引导(FSG),在扩散早期阶段优先解决长区间子问题,并增加迭代次数。在多种数据集和模型架构上的大量实验验证了FSG在图像质量和计算效率上均优于当前最优方法。本工作为条件引导提供了新视角,解锁了自适应设计的潜力。
原文摘要 · Abstract (English)
Classifier-Free Guidance (CFG) is an essential component of text-to-image diffusion models, and understanding and advancing its operational mechanisms remains a central focus of research. Existing approaches stem from divergent theoretical interpretations, thereby limiting the design space and obscuring key design choices. To address this, we propose a unified perspective that reframes conditional guidance as fixed point iterations, seeking to identify a golden path where latents produce consistent outputs under both conditional and unconditional generation. We demonstrate that CFG and its variants constitute a special case of single-step short-interval iteration, which is theoretically proven to exhibit inefficiency. To this end, we introduce Foresight Guidance (FSG), which prioritizes solving longer-interval subproblems in early diffusion stages with increased iterations. Extensive experiments across diverse datasets and model architectures validate the superiority of FSG over state-of-the-art methods in both image quality and computational efficiency. Our work offers novel perspectives for conditional guidance and unlocks the potential of adaptive design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。