无需额外模型,让扩散模型自我纠错生成更高质量图像
In-situ Autoguidance: Eliciting Self-Correction in Diffusion Models
- 用随机前向传播实时生成低质预测,实现自引导
- 零成本提升图像质量与提示对齐度,且保持多样性
- 适合追求高效生成的科研与应用开发者
高质量、多样化且与提示对齐的图像生成是扩散模型的核心目标。主流的无分类器引导(CFG)方法虽提升了质量和对齐度,却牺牲了多样性,导致二者纠缠。近期工作通过引入一个独立训练的劣化模型实现解耦,但需额外模型带来显著开销。本文提出「原位自引导」(In-situ Autoguidance),无需任何辅助组件,直接从模型自身动态生成劣化预测,将引导机制重构为推理时的自我修正。实验表明,该零成本方法不仅可行,更建立了一项新的高效引导基准,证明无需外部模型即可实现自引导优势。
原文摘要 · Abstract (English)
The generation of high-quality, diverse, and prompt-aligned images is a central goal in image-generating diffusion models. The popular classifier-free guidance (CFG) approach improves quality and alignment at the cost of reduced variation, creating an inherent entanglement of these effects. Recent work has successfully disentangled these properties by guiding a model with a separately trained, inferior counterpart; however, this solution introduces the considerable overhead of requiring an auxiliary model. We challenge this prerequisite by introducing In-situ Autoguidance, a method that elicits guidance from the model itself without any auxiliary components. Our approach dynamically generates an inferior prediction on the fly using a stochastic forward pass, reframing guidance as a form of inference-time self-correction. We demonstrate that this zero-cost approach is not only viable but also establishes a powerful new baseline for cost-efficient guidance, proving that the benefits of self-guidance can be achieved without external models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。