用动态偏好优化扩散采样器,提升低步数下的图像质量。
D2PO: Optimizing Diffusion Samplers via Dynamic Preference
- 将采样优化转为偏好学习,基于能量模型建模采样策略。
- 在低NFE条件下显著提升纹理细节与整体感知质量。
- 适合追求高质量生成结果的低资源生成应用。
我们提出D2PO(动态直接偏好优化),一种针对扩散采样器的时间步调度和无分类器引导(CFG)权重进行优化的原理性框架。现有学生-教师回归框架存在根本局限:低NFE学生采样器模仿高NFE教师,常牺牲高频纹理保真度而保留粗粒度结构,导致采样器与感知质量错位。D2PO通过将采样器优化重构为偏好对齐问题,利用直接偏好优化(DPO)框架解决该问题。为使DPO适用于扩散采样器,我们将采样策略建模为能量模型(EBM),将偏好比较转化为可计算的能量差。我们进一步提出一种直接源自预训练得分网络的新能量形式,可在扰动空间中联合捕捉结构一致性与细粒度细节。此外,引入动态偏好机制,随着采样策略学习,所用偏好样本逐步提升。该自增强机制取代静态教师监督,实现迭代式、偏好驱动的精炼过程,提供更强的对齐信号。大量实验表明,D2PO更忠实地对齐扩散采样器与感知质量,充分释放高质量教师的潜力,在低NFE约束下持续优于传统回归基线。
原文摘要 · Abstract (English)
We propose D2PO (Dynamic Direct Preference Optimization), a principled framework for optimizing diffusion sampling policies with respect to timestep schedules and classifier-free guidance (CFG) weights. Our work is motivated by a fundamental limitation of existing student-teacher regression frameworks; low-NFE student samplers are trained to mimic high-NFEteachers, often sacrificing high-frequency texture fidelity while preserving coarse global structures, thereby misaligning the sampler with perceptual quality. D2PO addresses this challenge by reformulating sampler optimization as a preference-based alignment problem, leveraging the Direct Preference Optimization (DPO) framework. To make DPO applicable to diffusion samplers, we model the sampling policy as an energy-based model (EBM), transforming preference comparisons into tractable energy differences. We further introduce a novel energy formulation derived directly from the pretrained score network, enabling preference evaluation in perturbed spaces that jointly capture structural consistency and fine-grained details. Moreover, we introduce dynamic preferences, where the preferred samples used for alignment progressively improve as the sampling policies are learned. This self-improving mechanism replaces rigid static teacher supervision with an iterative, preference-guided refinement process, providing progressively stronger alignment signals. Extensive experiments demonstrate that D2PO aligns diffusion samplers with perceptual quality more faithfully, unlocking the full potential of high-quality teachers and consistently outperforming conventional regression-based schedulers under low-NFE constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。