arXiv:2605.25123cs.LGcs.AI2026-05

提出新方法提升扩散模型推理时对齐效率,减少采样误差。

Inference-Time Alignment of Diffusion Models via Trust-Region Iterative Twisted Sequential Monte Carlo

论文配图:Inference-Time Alignment of Diffusion Models via Trust-Region Iterative Twisted Sequential Monte Carlo
图 1 · 摘自论文原文
  • 用信任域框架迭代优化扭曲函数,引导采样路径更高效。
  • 在相同计算预算下,文本生成与图文生成任务性能更优。
  • 适合需要高精度生成且不能修改模型权重的场景。

我们研究基于扩散模型的推理时对齐问题,旨在不更新模型权重的前提下,引导基模型生成高回报输出。现有基于序列蒙特卡洛(SMC)的方法虽能合理逼近奖励倾斜的目标分布,但其提议仍高度依赖基采样器。由于奖励信息主要通过粒子重加权和重采样传播,这些方法需大量粒子,易出现权重退化和高方差估计。为降低方差、提升粒子效率,可迭代学习提供前瞻指导的扭曲函数,如扭曲SMC。然而,现有可学习扭曲方法主要针对经典序列推断,在高维状态空间及终端噪声或黑箱奖励下的扩散模型对齐中易不稳定。本文提出信任域迭代扭曲序列蒙特卡洛(TRI-TSMC),一种用于SMC推理时对齐的学习扭曲函数框架。每次迭代在路径空间中进行精确的KL约束更新,通过温度重要性重加权获得闭式解,并通过加权最大似然将目标投影回参数化扭曲族。理论上,我们形式化了最优扭曲函数的价值函数解释,证明其可实现零方差采样;并证明信任域更新沿援助路径逼近目标分布,加权最大似然更新为前向KL投影,路径可降低残余重要性权重方差。实验表明,TRI-TSMC在离散扩散文本生成和文本到图像生成任务中,于匹配推理预算下显著提升主对齐目标性能。

原文摘要 · Abstract (English)

We study inference-time alignment for diffusion-based generative models, aiming to steer a base model toward high-reward outputs without updating its weights. Recent Sequential Monte Carlo (SMC)-based steering methods approximate reward-tilted target distributions in a principled way, but their proposals remain largely tied to the base sampler. Since reward information is mainly used after propagation through particle reweighting and resampling, these methods can require large particle budgets and suffer from weight degeneracy and high-variance estimates. One way to reduce variance and improve particle efficiency is to iteratively learn twisting functions that provide look-ahead guidance, as in twisted SMC. However, existing learnable twisting methods are developed mainly for classical sequential inference and can be unstable when applied to diffusion-based alignment with high-dimensional state spaces and terminal, noisy, or black-box rewards. We propose Trust-Region Iterative Twisted Sequential Monte Carlo (TRI-TSMC), a trust-region framework for learning twisting functions in SMC-based inference-time alignment. Each iteration computes an exact KL-constrained update in path space, which admits a closed-form solution by tempered importance reweighting, and projects this target back to the parameterized twisted family by weighted maximum likelihood. Theoretically, we formalize the value-function interpretation of the optimal twisting function and show that it yields a zero-variance sampler. We prove that the trust-region update follows an escort path toward the target distribution, that the weighted maximum-likelihood update is a forward-KL projection, and that the path reduces residual importance-weight variance. Empirically, TRI-TSMC improves primary alignment objectives on discrete diffusion text generation and text-to-image generation under matched inference-time budgets.

扩散模型推理对齐序列蒙特卡洛生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。