arXiv:2602.01591cs.CV2026-02被引 4

用分步优势优化提升流模型生成速度与偏好对齐效果

Faster and Better Alignment for Flow Matching Models via Step-aware Advantages

  • 通过温度退火采样引入动态噪声,增强生成多样性
  • 分步奖励机制让少步生成图像更贴合人类偏好
  • 无需可微奖励函数,适合高效文本到图像生成

近期基于强化学习的流匹配模型在少步文本到图像生成中显著提升了与人类偏好的对齐效果。然而,现有方法依赖大量去噪步骤,且奖励信号稀疏不精确,导致对齐不佳。为此,我们提出温度退火少步采样与组相对策略优化框架(TAFS-GRPO),通过迭代注入随时间变化的自适应噪声,对单步纯净预测进行扰动,在保持图像语义完整性的前提下引入随机性。其分步优势整合机制结合了GRPO与温度退火采样,无需可微奖励函数,提供密集且步骤特异的奖励信号,实现稳定策略优化。大量实验表明,TAFS-GRPO在少步生成中表现优异,显著提升生成图像与人类偏好的对齐度。代码与模型将公开以促进后续研究。

原文摘要 · Abstract (English)

Recent advances in flow matching models, particularly with reinforcement learning (RL), have significantly enhanced human preference alignment in few-step text-to-image generators. However, existing RL-based approaches for flow matching models typically rely on numerous denoising steps, while suffering from sparse and imprecise reward signals that often lead to suboptimal alignment. To address these limitations, we propose Temperature-Annealed Few-step Sampling with Group Relative Policy Optimization (TAFS-GRPO), a novel framework for training flow matching text-to-image models into efficient few-step generators well aligned with human preferences. Our method iteratively injects adaptive time-dependent noise into one-step clean predictions. By repeatedly annealing the model's sampled outputs, it introduces stochasticity into the sampling process while preserving the semantic integrity of each generated image. Moreover, its step-aware advantage integration mechanism combines GRPO with temperature-annealed sampling to eliminate the need for a differentiable reward function and provide dense, step-specific rewards for stable policy optimization. Extensive experiments demonstrate that TAFS-GRPO achieves strong performance in few-step text-to-image generation and significantly improves the alignment of generated images with human preferences. The code and models of this work will be available to facilitate further research.

文本到图像流匹配强化学习少步生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。