arXiv:2607.00535cs.LGcs.AI2026-07被引 1

让确定性生成模型通过强化学习优化采样路径。

Flow-Map GRPO: Reinforcement Learning for Few-Step Flow-Map Generators via Anchored Stochastic Composition

论文配图:Flow-Map GRPO: Reinforcement Learning for Few-Step Flow-Map Generators via Anchored Stochastic Composition
图 1 · 摘自论文原文
  • 提出锚点条件重采样机制,保留原路径的同时引入随机性
  • 在文生图模型上提升感知质量与任务表现,超越原始模型
  • 无需修改结构或重新训练,即可实现后置强化学习优化

少步流图生成器(如一致性模型和MeanFlow)通过直接学习噪声到数据的长程传输映射加速采样。然而这些模型通常为确定性,难以用需随机轨迹和明确定义似然比的强化学习方法进行后训练优化。现有基于SDE的随机化技术适用于基于速度的采样器及无穷小或精细离散化转移,不适用于长程流图。本文提出Flow-Map GRPO,一种面向确定性少步流图生成器的在线强化学习后训练框架。核心是锚点式随机流图组合(ASFMC),通过锚点条件重采样引入随机性,同时保持原流图的边缘概率路径。推导了单时间与双时间流图参数化的GRPO目标。在基于FLUX的少步文生图生成器(包括MeanFlow和sCM)上实验表明,Flow-Map GRPO在奖励、感知和任务级评估指标上均优于预训练确定性模型。结果证明,无需修改原模型参数化或将其重训练为原生随机模型,即可有效对齐确定性少步流图生成器与强化学习后训练。

原文摘要 · Abstract (English)

Few-step flow-map generators, such as consistency models and MeanFlow, accelerate sampling by directly learning long-range transport maps between noise and data. However, these models are typically deterministic, which makes them difficult to optimize with reinforcement learning (RL) post-training methods that require stochastic trajectories and well-defined likelihood ratios. Existing SDE-based stochasticization techniques are designed for velocity-based samplers with infinitesimal or finely discretized transitions, and therefore do not directly apply to long-range flow maps. In this work, we propose Flow-Map GRPO, an online RL post-training framework for deterministic few-step flow-map generators. The key component is Anchored Stochastic Flow Map Composition (ASFMC), a path-preserving stochasticization mechanism that introduces randomness through anchor-based conditional resampling while preserving the original marginal probability path of the deterministic flow map. We derive GRPO objectives for both single-time and two-time flow-map parameterizations. Experiments on few-step FLUX-based text-to-image generators, including MeanFlow and sCM, show that Flow-Map GRPO improves pretrained deterministic flow-map models across reward-based, perceptual, and task-level evaluation metrics. Our results demonstrate that deterministic few-step flow-map generators can be effectively aligned with RL post-training without modifying their original model parameterization or retraining them as native stochastic models.

生成模型强化学习流图生成后训练优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。