arXiv:2608.04887cs.CV2026-08

让扩散模型学生超越老师,通过优化内部演化路径提升生成能力

STEP-OPD: Rethinking Output Targets and Internal Dynamics in On-Policy Distillation for Diffusion Models

论文配图:STEP-OPD: Rethinking Output Targets and Internal Dynamics in On-Policy Distillation for Diffusion Models
图 1 · 摘自论文原文
  • 用教师与基础模型的输出速度差作为新学习方向,扩展学生目标
  • 在组合对齐、文本渲染等任务上,GenEval得分从0.927提升至0.961
  • 适合需要统一多任务生成能力的研究者,尤其关注内部表征演化

在策略蒸馏(OPD)中,现有方法主要使学生匹配教师的输出速度,将教师视为优化上限。但仅依赖输出监督会弱化学生各层表征演化的约束,阻碍能力的逐步迁移。本文提出STEP-OPD,将学生学习目标延伸至教师之外,并显式约束其内部表示演化过程。不把教师当作最终目标,而是利用每个任务专属教师与共享基础模型之间的速度差作为进一步学习的方向,将该差值的缩放版本加到教师速度上。同时,对齐学生与教师在各网络块间表征变化的方向与幅度,使学生学会表征如何逐层演变。在组合对齐、文本渲染和人类偏好测试中,本方法持续优于标准OPD。GenEval得分由DiffusionOPD的0.927提升至0.961,同时在OCR及所有偏好指标上均获改善。最终统一的学生模型在三类能力上均超越对应单任务教师,证明输出外推实现超越教师的学习;表征变化对齐则为内部转换提供了互补指导。

原文摘要 · Abstract (English)

On-policy distillation (OPD) has become an effective approach for consolidating multiple task-specialized image generation models into a single student. However, existing OPD methods optimize the student mainly to match the teacher's output velocity, making the teacher the upper limit of the optimization objective. While output-level supervision alone leaves the student's blockwise representation evolution underconstrained, which weakens the transfer of capabilities that must be progressively developed across layers. We propose STEP-OPD, an on-policy distillation framework for image generation that extends the student's learning target beyond the teacher and introduces explicit constraints on its internal representation evolution. Instead of treating the teacher as the final target, we use the velocity difference between each task-specific teacher and the shared base model as a direction for further learning and add a scaled version of this difference to the teacher velocity. In addition, we align the direction and magnitude of representation changes between the student and teacher, enabling the student to learn how representations are progressively transformed across network blocks. Experiments on compositional alignment, text rendering, and human preference show that our method consistently improves Standard OPD methods. In particular, it increases the GenEval score of DiffusionOPD from 0.927 to 0.961, while also improving OCR and all preference-based metrics. The resulting unified student surpasses the corresponding single-task teachers across all three capability groups, showing that output extrapolation enables beyond-teacher learning. And representation change alignment provides complementary guidance for the student's internal transformations.

扩散模型知识蒸馏表征演化生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。