arXiv:2606.30376cs.LGcs.CV2026-06被引 1

提出无需引导的流模型优化方法,加速生成并提升对齐效果。

FlowAWR: Online Adaptive Flow Reinforcement via Advantage-Weighted Rectification

论文配图:FlowAWR: Online Adaptive Flow Reinforcement via Advantage-Weighted Rectification
图 1 · 摘自论文原文
  • 将生成策略优化转为监督回归,直接学习最优速度场。
  • 在SD3.5-Medium上24.12分PickScore仅需1.2k步,提速2~5倍。
  • 无需分类器自由引导,支持多奖励约束下稳定生成。

通过在线强化学习对连续空间中的生成流模型进行对齐,受限于难以计算的轨迹似然。现有基于密度近似的策略梯度方法依赖随机SDE采样器构建可处理的转移核,引入训练与推理不一致问题,并需要分类器自由引导(CFG)。而像DiffusionNFT这样的隐式框架虽直接优化前向过程速度场,但其启发式的固定幅度修正无法根据组内质量相对强度调整优化力度。本文提出流优势加权修正(FlowAWR),将连续生成策略优化重构为向理论最优速度场进行监督回归。从带KL约束的奖励最大化最优策略出发,推导出具有幅度感知和优势加权修正形式的最优速度场,实现无SDE、无CFG的生成优化。在SD3.5-Medium上的对比实验显示,FlowAWR在保持对齐性能的同时,收敛速度比DiffusionNFT快2至5倍(如1.2k步达到24.12的PickScore,而DiffusionNFT需2.0k步达23.82,FlowGRPO则需超过4k步达23.50)。在多奖励约束下,FlowAWR仍能维持生成质量,满足结构规则并保持稳定的域外表现。

原文摘要 · Abstract (English)

Aligning generative flow models on continuous spaces via online reinforcement learning is constrained by intractable trajectory likelihoods. Existing density-approximated policy gradient methods rely on stochastic SDE samplers to construct tractable transition kernels, which introduce training-inference inconsistencies and necessitates Classifier-Free Guidance (CFG). While implicit frameworks such as DiffusionNFT directly optimize forward-process velocity fields, its heuristic fixed-magnitude corrections prevent optimization strength from relative intra-group quality. We propose \textit{Flow Advantage-Weighted Rectification} (\textbf{FlowAWR}), a paradigm that recasts continuous generative policy optimization as supervised regression toward a theoretically optimal velocity field. Starting from the optimal policy of a KL-constrained reward maximization, FlowAWR derives the optimal velocity field that admits a magnitude-aware, advantage-weighted rectification form, yielding SDE-free optimization and CFG-free generation. In comparative evaluations on SD3.5-Medium, FlowAWR achieves improved alignment performance alongside a 2$\times$ to 5$\times$ convergence acceleration over DiffusionNFT (e.g., reaching a 24.12 PickScore in 1.2k steps, versus 23.82 in 2.0k steps for DiffusionNFT and 23.50 in $>$4k steps for FlowGRPO). Under multi-reward constraints, FlowAWR sustains generation quality, satisfying structural rules while maintaining stable out-of-domain performance.

生成模型强化学习流模型无引导生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。