arXiv:2602.16198cs.LG2026-02被引 3

无需训练即可高效适配扩散模型,支持任意奖励函数。

Training-Free Adaptation of Diffusion Models via Doob's $h$-Transform

  • 基于杜布变换构建采样过程动态修正机制
  • 在D4RL离线强化学习基准上性能超越现有方法
  • 无需训练、计算高效,适用于不可导奖励

适配方法是释放预训练扩散模型在多样化应用中潜力的关键。现有方法常将适配目标抽象为奖励函数,并引导扩散模型生成高奖励样本,但往往需要额外训练导致计算开销大,或对奖励函数的可微性有严格要求。此外,尽管实证表现良好,理论依据和保证通常缺失。本文提出DOIT(Doob-Oriented Inference-time Transformation),一种无需训练且计算高效的适配方法,适用于任意非可导奖励。其核心框架为测度传输形式,旨在将预训练生成分布转移到高奖励目标分布。我们利用杜布的$h$-变换实现该传输,对扩散采样过程引入动态修正,可在不修改预训练模型的前提下实现高效模拟计算。理论上,我们通过刻画动态杜布修正的近似误差,建立了以高概率收敛至目标高奖励分布的保证。实验表明,在D4RL离线强化学习基准上,该方法持续优于当前最优基线,同时保持采样效率。

原文摘要 · Abstract (English)

Adaptation methods have been a workhorse for unlocking the transformative power of pre-trained diffusion models in diverse applications. Existing approaches often abstract adaptation objectives as a reward function and steer diffusion models to generate high-reward samples. However, these approaches can incur high computational overhead due to additional training, or rely on stringent assumptions on the reward such as differentiability. Moreover, despite their empirical success, theoretical justification and guarantees are seldom established. In this paper, we propose DOIT (Doob-Oriented Inference-time Transformation), a training-free and computationally efficient adaptation method that applies to generic, non-differentiable rewards. The key framework underlying our method is a measure transport formulation that seeks to transport the pre-trained generative distribution to a high-reward target distribution. We leverage Doob's $h$-transform to realize this transport, which induces a dynamic correction to the diffusion sampling process and enables efficient simulation-based computation without modifying the pre-trained model. Theoretically, we establish a high probability convergence guarantee to the target high-reward distribution via characterizing the approximation error in the dynamic Doob's correction. Empirically, on D4RL offline RL benchmarks, our method consistently outperforms state-of-the-art baselines while preserving sampling efficiency.

扩散模型无训练适配强化学习杜布变换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。