无需训练即可提升桥接模型的生成质量,通过先验引导实现
GuidedBridge: Training-freely Improving Bridge Models with Prior Guidance

- 利用未见先验与已见先验对比,通过缩放因子增强先验利用
- 提出频域调制引导(FMPG),按频率带调节引导强度,适配生成动态
- 构建CFG-FMPG级联框架,兼顾图像修复效率与生成质量
指导方法如无分类器引导(CFG)和自动引导(AG)推动了扩散模型中从噪声到数据的生成。最近,桥接模型引入了数据到数据的生成流程,可利用清晰的先验信息。受此前通过去噪结果差异制造质量差异的启发,本文提出一种无需训练的桥接引导方法——先验引导(PG)。具体地,引入一个预训练阶段未见过的弱先验,阻碍先验利用并降低去噪效果;随后与可见先验对比,通过缩放因子突出并强化先验利用。进一步分析桥接过程中的先验利用机制,设计频率调制先验引导(FMPG),将引导尺度适配低频与高频成分,符合桥接生成动态。针对图像修复任务,提出级联框架CFG-FMPG:先用CFG生成噪声隐表示,再以该表示为生成先验,通过FMPG进行优化,充分发挥二者优势而不影响推理效率。实验表明,所提PG方法在多种图像转换任务中均一致提升预训练桥接模型性能。
原文摘要 · Abstract (English)
Guidance methods, such as classifier-free guidance (CFG) and auto-guidance (AG), have advanced noise-to-data generation in diffusion models. Recently, bridge models have introduced a data-to-data generative process that can exploit an instructive clean prior. In this work, inspired by previous methods creating quality difference between denoising results as guidance, we propose a training-free bridge guidance method, termed Prior Guidance (PG). Specifically, we introduce a weak prior, which is unseen during bridge pre-training, hindering prior exploitation and thereby degrading denoising result. Then, we contrast it with the seen prior to highlight and enhance prior exploitation via a scaling factor. Moreover, we analyze the underlying mechanism of prior exploitation in the bridge process and design frequency-modulated prior guidance (FMPG), which tailors the guidance scale to low- and high-frequency bands coherent with bridge generative dynamics. To address prior exploitation in image in-painting, we develop a cascaded framework, CFG-FMPG, which first generates a noisy hidden representation via CFG and then exploits it as a generative prior with FMPG, fulfilling their complementary strengths without compromising inference efficiency. Experiments demonstrate that our PG methods consistently improve pre-trained bridge models across diverse image translation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。