arXiv:2510.02654cs.CV2025-10被引 10

让流匹配模型学会聪明地加噪,提升图像生成质量与对齐性。

Smart-GRPO: Smartly Sampling Noise for Efficient RL of Flow-Matching Models

  • 通过迭代搜索优化噪声分布,智能选择高奖励方向的扰动。
  • 相比基线方法,奖励优化与图像质量均显著提升。
  • 适合希望用强化学习改进文本到图像生成的科研与工程人员。

流匹配技术近期推动了高质量文本到图像生成的发展。然而,流匹配模型的确定性特性使其难以直接用于强化学习,而强化学习是提升图像质量与人类对齐性的关键工具。以往工作通过在潜在空间中加入随机噪声引入随机性,但此类扰动效率低且不稳定。本文提出 Smart-GRPO,首个针对流匹配模型强化学习的噪声优化方法。Smart-GRPO 采用迭代搜索策略:解码候选扰动,用奖励函数评估,并引导噪声分布向高奖励区域收敛。实验表明,相较于基线方法,Smart-GRPO 在奖励优化与视觉质量上均有显著提升。结果表明,该方法为流匹配框架中的强化学习提供了可行路径,弥合了高效训练与人类对齐生成之间的差距。

原文摘要 · Abstract (English)

Recent advancements in flow-matching have enabled high-quality text-to-image generation. However, the deterministic nature of flow-matching models makes them poorly suited for reinforcement learning, a key tool for improving image quality and human alignment. Prior work has introduced stochasticity by perturbing latents with random noise, but such perturbations are inefficient and unstable. We propose Smart-GRPO, the first method to optimize noise perturbations for reinforcement learning in flow-matching models. Smart-GRPO employs an iterative search strategy that decodes candidate perturbations, evaluates them with a reward function, and refines the noise distribution toward higher-reward regions. Experiments demonstrate that Smart-GRPO improves both reward optimization and visual quality compared to baseline methods. Our results suggest a practical path toward reinforcement learning in flow-matching frameworks, bridging the gap between efficient training and human-aligned generation.

强化学习流匹配图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。