让扩散模型采样时间自适应调整,提升生成质量。
ART for Diffusion Sampling: Continuous-Time Control and Actor-Critic Learning

- 将采样时间控制建模为连续时间强化学习问题,自动优化时间步分配。
- 在图像生成等任务中,相同计算量下显著提升样本质量。
- 学到的时间调度可跨数据集、模型和预算直接迁移使用。
我们研究基于得分的扩散采样中的时间步分配问题,其中学习到的逆向动态在有限网格上离散化。均匀和人工设计的时间表是常用选择,但依赖固定规则,可能次优。为此,我们提出自适应重参数化时间(ART),一种连续时间控制框架,将采样时钟速度视为控制变量,使学习后的时钟上均匀网格在原始扩散时间中诱导出自适应时间步。基于一阶欧拉误差近似,ART 提供了沿采样轨迹分配时间步的合理目标。为求解此确定性控制问题,我们引入 ART-RL,一种带有高斯策略的辅助随机化形式,将调度学习转化为连续时间强化学习问题。我们证明,在最优解层面,随机化的 ART-RL 与 ART 等价,其最优高斯策略通过均值恢复最优 ART 时间扭曲率。我们进一步建立策略评估与改进刻画,并推导出可用于实现演员-批评家更新的轨迹级矩恒等式。在从受控低维设置到图像生成的实验中,仅更改时间步网格,ART-RL 可无缝集成到现有扩散采样器中,始终在相同预算下超越强基线调度,且不改变采样流程其余部分。所学调度展现出广泛泛化能力,无需重新训练即可跨采样预算、数据集、求解器、流水线和表示空间迁移。
原文摘要 · Abstract (English)
We study timestep allocation for score-based diffusion sampling, where a learned reverse-time dynamics is discretized on a finite grid. Uniform and hand-crafted schedules are standard choices, but they rely on fixed prescriptions and can therefore be suboptimal. To address this limitation, we propose Adaptive Reparameterized Time (ART), a continuous-time control formulation that learns a time change by treating the speed of the sampling clock as the control, so that a uniform grid on the learned clock induces adaptive timesteps in the original diffusion time. Based on a leading-order Euler error surrogate, ART provides a principled objective for allocating timesteps along the sampling trajectory. To solve this deterministic control problem, we introduce ART-RL, an auxiliary randomized formulation with Gaussian policies that turns schedule learning into a continuous-time reinforcement learning problem. We prove that the randomized ART-RL formulation is equivalent to ART at the optimizer level, in the sense that its optimal Gaussian policy recovers the optimal ART time-warping rate through its mean. We further establish policy evaluation and policy improvement characterizations and derive trajectory-based moment identities that yield implementable actor--critic updates for learning the schedule. Across experiments ranging from controlled low-dimensional settings to image generation, ART-RL can be plugged into existing diffusion samplers by changing only the timestep grid, consistently improving sample quality over strong baseline schedules at matched budgets while leaving the rest of the sampling pipeline unchanged. The learned schedules also exhibit broad generalization, transferring without retraining across sampling budgets, datasets, solvers, pipelines, and representation spaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。