用分层变分策略降低扩散模型推理成本,提速超5倍且质量更好。
Hierarchical Variational Policies for Reward-Guided Diffusion

- 将测试时优化转化为分层变分模型,用轻量策略实现快速控制
- 4倍超分任务中推理速度提升5倍以上,感知质量更优
- 适合需要高效生成的图像逆问题场景,如医学成像
将预训练扩散模型适配至下游目标(如逆问题)通常需要昂贵的测试时引导或优化。本文提出一种原则性框架,在显著降低推理成本的同时生成高质量奖励对齐样本。方法将测试时适应建模为分层变分模型,将控制机制压缩为一个轻量但表达力强的随机策略。该框架天然支持少步扩散采样:大步长实现快速推理,而学习到的策略通过每步结构化控制维持样本质量。最终的全摊销采样器在质量和速度间取得优异权衡,性能匹配甚至超越近期测试时扩展基线,但计算开销大幅降低。例如,在4倍超分辨率任务中,本方法比最优基线推理更快5倍以上,且感知质量更优。此外,还将方法扩展至半摊销范式,结合廉价摊销提议与有限测试时优化,在多个挑战性逆问题上达到当前最优感知质量。
原文摘要 · Abstract (English)
Adapting pretrained diffusion models to downstream objectives such as inverse problems often requires expensive test-time guidance or optimization. We propose a principled framework for generating high-quality reward-aligned samples at substantially reduced inference cost. Our approach formulates test-time adaptation as a hierarchical variational model, where control is amortized into a lightweight yet expressive stochastic policy. This formulation naturally supports few-step diffusion sampling: large step sizes enable fast inference, while the learned policy maintains sample quality by providing structured per-step control. The resulting fully amortized sampler achieves a strong quality--speed tradeoff, matching or exceeding recent test-time scaling baselines while requiring significantly less compute. For example, on 4x super-resolution, our method achieves better perceptual quality with more than 5x faster inference compared to the best-performing baseline. We further extend our approach to a semi-amortized regime that combines cheap amortized proposals with limited test-time optimization, achieving state-of-the-art perceptual quality across several challenging inverse problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。