arXiv:2604.14910cs.CV2026-04被引 1

让少步生成模型超越教师,通过奖励感知引导优化方向。

Reward-Aware Trajectory Shaping for Few-step Visual Generation

论文配图:Reward-Aware Trajectory Shaping for Few-step Visual Generation
图 1 · 摘自论文原文
  • 用奖励感知门控动态调节教师指导,避免盲目模仿。
  • 在4步生成中质量超越多步教师,显著缩小性能差距。
  • 无需额外计算开销,适合高效图像生成应用。

在极少数采样步骤内实现高质量生成是生成建模的核心目标。现有方法主要依赖蒸馏框架,将多步去噪过程压缩为少步生成器,但这类方法强制学生模仿更强的多步教师,使教师成为性能上限。本文提出奖励感知轨迹塑造(RATS),通过在关键去噪阶段对齐师生潜在轨迹,并引入奖励感知门控,根据相对奖励表现自适应调节教师引导强度。当教师奖励更高时强化轨迹塑造,当学生匹配或超越教师时则放松约束,从而实现持续的奖励驱动优化。RATS无缝融合轨迹蒸馏、奖励感知门控与偏好对齐,有效传递高步数生成器中的偏好相关知识,且测试时无额外计算开销。实验表明,RATS显著提升了少步生成的效率-质量权衡,大幅缩小了少步学生与强多步生成器之间的差距。

原文摘要 · Abstract (English)

Achieving high-fidelity generation in extremely few sampling steps has long been a central goal of generative modeling. Existing approaches largely rely on distillation-based frameworks to compress the original multi-step denoising process into a few-step generator. However, such methods inherently constrain the student to imitate a stronger multi-step teacher, imposing the teacher as an upper bound on student performance. We argue that introducing \textbf{preference alignment awareness} enables the student to optimize toward reward-preferred generation quality, potentially surpassing the teacher instead of being restricted to rigid teacher imitation. To this end, we propose \textbf{Reward-Aware Trajectory Shaping (RATS)}, a lightweight framework for preference-aligned few-step generation. Specifically, teacher and student latent trajectories are aligned at key denoising stages through horizon matching, while a \textbf{reward-aware gate} is introduced to adaptively regulate teacher guidance based on their relative reward performance. Trajectory shaping is strengthened when the teacher achieves higher rewards, and relaxed when the student matches or surpasses the teacher, thereby enabling continued reward-driven improvement. By seamlessly integrating trajectory distillation, reward-aware gating, and preference alignment, RATS effectively transfers preference-relevant knowledge from high-step generators without incurring additional test-time computational overhead. Experimental results demonstrate that RATS substantially improves the efficiency--quality trade-off in few-step visual generation, significantly narrowing the gap between few-step students and stronger multi-step generators.

少步生成奖励对齐扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。