通过自适应终止采样,提升扩散模型的高效与精准生成。
PAST: Prompt-Adaptive Sampling Termination for Efficient Diffusion Model

- 根据去噪进度和提示难度动态调整训练时长。
- 计算效率提升66.7%,偏好优化质量提高29.5%。
- 适合需要高效微调扩散模型的研究者使用。
尽管扩散模型在文本到图像任务中取得显著进展,但在直接优化下游目标时仍存在局限。虽然强化学习(RL)可实现定向优化,但现有方法普遍受限于低效微调和稀疏奖励。为此,我们提出PAST,该方法通过联合感知去噪进度与提示难度,提供差异化奖励并自适应调节训练周期。我们设计了内在奖励机制以弥补外在奖励的稀疏性,引导模型更高效地脱离噪声模式。进一步地,PAST动态监测去噪完成度与图像结构和提示语义间的对齐程度,当两者均满足生成要求时,系统自适应终止训练,实现基于提示难度与当前生成过程的合理周期分配。最后,基于预测残差噪声水平,构建双重自适应协调机制,平衡外在与内在奖励,同时兼顾探索与收敛。实验表明,PAST将现有强化学习微调方法的计算效率提升高达66.7%,并通过双重自适应调节机制使偏好优化质量提升29.5%。
原文摘要 · Abstract (English)
While diffusion models have made significant progress in text-to-image tasks, they still exhibit limitations when directly optimizing downstream objectives. Although Reinforcement Learning (RL) enables targeted optimization, existing methods are generally constrained by low-efficiency fine-tuning and sparse rewards. To address these challenges, we propose PAST, which provides differentiated rewards while adaptively regulating training episode length by jointly perceiving denoising progress and prompt difficulty. Specifically, we design an intrinsic reward paradigm to compensate for sparse extrinsic rewards and guide the model to explore paths that diverge more efficiently from noise patterns. We further provide theoretical justification for intrinsic rewards. Then, PAST dynamically monitors denoising completion and semantic alignment between image structures and prompt semantics. When both metrics satisfy generation requirements, the system adaptively terminates training. This enables appropriate allocation of episode lengths based on prompt difficulty and the current generation process. Finally, based on the predicted residual noise level, we establish a dual adaptive coordination mechanism. Specifically, it not only balances the extrinsic and intrinsic rewards but also balances the exploration and convergence. Experimental results demonstrate that PAST enhances computational efficiency of existing RL fine-tuning methods by up to 66.7%, while improving preference optimization quality by up to 29.5% through its dual adaptive regulation mechanism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。