让流模型适应新任务更稳定高效,避免干扰与高方差问题。
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models

- 用退化参考增强对比,实现可外推的稳定训练
- 在单/多教师设置下性能优于传统方法和专用教师
- 适合需要快速适配新任务的生成模型应用
流匹配模型已成为图像生成的主流方法,但其在下游场景中的适配通常依赖于后训练,可能引发任务间优化目标冲突。强化学习虽能直接优化特定任务奖励,但轨迹级优化易导致高方差梯度与跨任务干扰。在线策略蒸馏(OPD)为学生轨迹提供密集且稳定的监督,但传统教师匹配仍基于模仿学习。本文提出DreOPD——一种用于流匹配模型的退化参考外推型在线策略蒸馏方法,融合两种范式。DreOPD将隐式奖励外推转化为闭式速度回归,实现稳定外推后训练。同时引入轻微退化的参考以增强教师-参考对比,明确外推方向。在单教师与多教师设置下的实验表明,DreOPD在平均性能上优于OPD与多任务强化学习基线,且在多数指标上超越专用教师。
原文摘要 · Abstract (English)
Flow-matching models are now a mainstream method to image generation, but its adaptation to diverse downstream scenarios typically relies on post-training, which may cause conflicts among task-specific optimization objectives. Reinforcement learning enables direct optimization of task-specific rewards beyond the original models, yet trajectory-level optimization may incur high-variance gradients and cross-task interference. On-policy distillation (OPD) offers dense and stable supervision on student rollouts, but conventional teacher matching remains imitation-based. We propose DreOPD, a Degraded-reference extrapolative OPD method for flow-matching models that bridges these two paradigms. Our DreOPD converts implicit reward extrapolation into closed-form velocity regression, enabling extrapolative post-training with the stability of OPD. It further uses a mildly degraded reference to strengthen the teacher-reference contrast, yielding a clearer extrapolation direction. Experiments on single- and multi-teacher settings show that DreOPD outperforms OPD and multi-task RL baselines in average performance, while surpassing specialized teachers on most metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。