arXiv:2607.15273cs.CVcs.LG2026-07

将强化学习引入平均速度生成模型,实现快速高效的内容生成。

MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators

论文配图:MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators
图 1 · 摘自论文原文
  • 构建瞬时速度预测器,衔接平均速度与强化学习优化
  • 4步采样即达84.33分(VBench),超越50步传统方法
  • 保留快速生成优势,适用于图像与视频生成任务

均值流生成器通过预测时间区间内的平均速度实现快速少步采样,极具效率优势。强化学习(RL)已成为对齐扩散与流模型与人类偏好及任务目标的强大工具。特别是,DiffusionNFT 提供了一种无需反向轨迹或似然估计的前向过程强化学习框架。然而,该方法在均值流上的应用仍不充分。DiffusionNFT 优化瞬时速度,而均值流依赖平均速度。为此,我们提出 MeanFlowNFT。受均值流恒等式启发,我们构建了一个诱导的瞬时速度预测器,并将其用于 DiffusionNFT 目标,使均值流的奖励优化可定义。采样仍基于平均速度,保持均值流的快速生成特性。我们进一步证明,MeanFlowNFT 继承了 DiffusionNFT 的严格策略改进保证。在图像与视频生成任务上,实验表明其持续优于基线。尤其在 SD3.5-M 上,8项指标中有6项超越当前最优,且仅用少数采样步数即可超过多步强化学习调优的扩散模型。例如,在 Wan 2.1 上,4步的 MeanFlowNFT 达到 84.33 的 VBench 分数,高于 50 步的 LongCat-Video RL(82.57)。

原文摘要 · Abstract (English)

MeanFlow generators achieve fast few-step sampling by predicting average velocities over time intervals, making them attractive for efficient generation. Reinforcement learning (RL) has become a powerful way to align diffusion and flow models with human preferences and task-specific objectives. In particular, DiffusionNFT offers an efficient forward-process RL framework that does not require reverse-process trajectories or likelihood estimation. However, applying such RL methods to MeanFlow remains underexplored. DiffusionNFT optimizes instantaneous velocities, whereas MeanFlow samples with average velocities. To bridge this gap, we introduce MeanFlowNFT. Inspired by the MeanFlow identity, which bridges average and instantaneous velocities, we construct an induced instantaneous-velocity predictor. We apply the DiffusionNFT objective to this predictor, making reward optimization well-defined for MeanFlow. Sampling remains based on the average velocity, preserving MeanFlow's fast few-step generation. We further prove that MeanFlowNFT inherits DiffusionNFT's strict policy-improvement guarantee. Experiments on image and video generation show that MeanFlowNFT consistently improves baselines. Moreover, it outperforms prior state-of-the-art RL-tuned few-step generators on most metrics ($6$ of $8$ on SD3.5-M), and can even surpass multi-step RL-tuned diffusion while using only a few sampling steps. For instance, on Wan 2.1, $4$-step MeanFlowNFT reaches a VBench score of $84.33$, surpassing $50$-step LongCat-Video RL ($82.57$).

生成模型强化学习快速采样视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。