揭示两种生成控制方法的统一本质,实现计算与精度的灵活权衡。
Greed is Good: A Unifying Perspective on Guided Generation
- 将后验引导视为端到端引导的贪婪策略,建立理论关联
- 提出插值方法,在计算成本与梯度精度间动态平衡
- 在图像逆问题与分子生成中验证有效性
无需训练的引导生成是广泛使用且强大的技术,可让用户对流/扩散模型的生成过程施加额外控制。目前针对基于梯度的引导,主要发展出两类方法:后验引导(通过目标预测模型将当前样本投影至目标分布)和端到端引导(在整个ODE求解过程中进行反向传播)。本文表明,这两类看似独立的方法实际上可通过将后验引导视为端到端引导的贪婪策略而统一。我们深入分析了两类方法与连续理想梯度之间的理论关系,并据此提出一种在两者间插值的方法,实现计算开销与引导梯度精度间的权衡。我们在多个逆图像问题及属性引导分子生成任务上验证了该方法的有效性。
原文摘要 · Abstract (English)
Training-free guided generation is a widely used and powerful technique that allows the end user to exert further control over the generative process of flow/diffusion models. Generally speaking, two families of techniques have emerged for solving this problem for gradient-based guidance: namely, posterior guidance (i.e., guidance via projecting the current sample to the target distribution via the target prediction model) and end-to-end guidance (i.e., guidance by performing backpropagation throughout the entire ODE solve). In this work, we show that these two seemingly separate families can actually be unified by looking at posterior guidance as a greedy strategy of end-to-end guidance. We explore the theoretical connections between these two families and provide an in-depth theoretical of these two techniques relative to the continuous ideal gradients. Motivated by this analysis we then show a method for interpolating between these two families enabling a trade-off between compute and accuracy of the guidance gradients. We then validate this work on several inverse image problems and property-guided molecular generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。