arXiv:2603.12648cs.CV2026-03被引 2

通过扩展条件空间实现多视角评估,提升文本生成图像模型对齐效果

From Sparse to Dense: Multi-View GRPO for Flow Models via Augmented Condition Space

  • 用语义相近的多组提示词生成多视角奖励信号
  • 在不重生成样本情况下提升对齐性能,优于当前最优方法
  • 适合需要精细控制生成结果的图像生成研究者

Group Relative Policy Optimization (GRPO) 已成为文本到图像(T2I)流模型中偏好对齐的强大框架。然而,我们发现标准范式——将一组生成样本与单个条件进行评估——因样本间关系探索不足,限制了对齐效果和性能上限。为解决这一稀疏单视角评估问题,本文提出 Multi-View GRPO(MV-GRPO),通过增强条件空间构建密集多视角奖励映射,以促进关系探索。具体而言,对于同一提示生成的一组样本,MV-GRPO 利用灵活的 Condition Enhancer 生成语义相邻但多样化的描述词。这些新提示支持多视角优势再估计,捕捉多样化语义属性,提供更丰富的优化信号。通过推导原始样本在新提示条件下的概率分布,可在无需昂贵样本重生成的情况下将其融入训练过程。大量实验表明,MV-GRPO 在对齐性能上显著优于现有最先进方法。

原文摘要 · Abstract (English)

Group Relative Policy Optimization (GRPO) has emerged as a powerful framework for preference alignment in text-to-image (T2I) flow models. However, we observe that the standard paradigm where evaluating a group of generated samples against a single condition suffers from insufficient exploration of inter-sample relationships, constraining both alignment efficacy and performance ceilings. To address this sparse single-view evaluation scheme, we propose Multi-View GRPO (MV-GRPO), a novel approach that enhances relationship exploration by augmenting the condition space to create a dense multi-view reward mapping. Specifically, for a group of samples generated from one prompt, MV-GRPO leverages a flexible Condition Enhancer to generate semantically adjacent yet diverse captions. These captions enable multi-view advantage re-estimation, capturing diverse semantic attributes and providing richer optimization signals. By deriving the probability distribution of the original samples conditioned on these new captions, we can incorporate them into the training process without costly sample regeneration. Extensive experiments demonstrate that MV-GRPO achieves superior alignment performance over state-of-the-art methods.

图像生成扩散模型强化学习对齐优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。