arXiv:2512.21514cs.CVcs.AI2025-12被引 14

解决图像生成中后期模式坍缩问题,提升视觉多样性。

DiverseGRPO: Mitigating Mode Collapse in Image Generation via Diversity-Aware GRPO

  • 引入分布级创意奖励,按语义分组动态分配探索奖励。
  • 实验显示质量相当下多样性提升13%~18%,突破原有平衡点。
  • 适合关注生成多样性与真实感的图像生成研究者使用。

强化学习(尤其是GRPO)通过组内图像相对比较显著提升了图像生成质量。然而,在训练后期,模型倾向于生成同质化输出,缺乏创意和视觉多样性,限制了应用范围。该问题可从奖励建模和生成动态两方面分析:传统GRPO依赖单样本质量作为奖励信号,导致模型收敛至少数高奖励生成模式,忽视整体分布多样性;同时,常规正则化未充分考虑早期去噪对多样性保持的关键作用,造成正则化预算错配,制约了质量与多样性的权衡。针对上述问题,本文从奖励建模与生成动态双重角度重新审视多样性退化机制。在奖励层面,提出基于语义分组的分布级创意奖励:通过对同一标题生成的样本进行谱聚类构建分布表示,并根据组大小自适应分配探索性奖励,鼓励发现新视觉模式。在生成层面,引入结构感知正则化,增强早期阶段约束以保留多样性,同时不损害奖励优化效率。实验表明,本方法在匹配质量评分下实现13%~18%的语义多样性提升,为基于GRPO的图像生成建立了新的质量-多样性帕累托前沿。

原文摘要 · Abstract (English)

Reinforcement learning (RL), particularly GRPO, improves image generation quality significantly by comparing the relative performance of images generated within the same group. However, in the later stages of training, the model tends to produce homogenized outputs, lacking creativity and visual diversity, which restricts its application scenarios. This issue can be analyzed from both reward modeling and generation dynamics perspectives. First, traditional GRPO relies on single-sample quality as the reward signal, driving the model to converge toward a few high-reward generation modes while neglecting distribution-level diversity. Second, conventional GRPO regularization neglects the dominant role of early-stage denoising in preserving diversity, causing a misaligned regularization budget that limits the achievable quality--diversity trade-off. Motivated by these insights, we revisit the diversity degradation problem from both reward modeling and generation dynamics. At the reward level, we propose a distributional creativity bonus based on semantic grouping. Specifically, we construct a distribution-level representation via spectral clustering over samples generated from the same caption, and adaptively allocate exploratory rewards according to group sizes to encourage the discovery of novel visual modes. At the generation level, we introduce a structure-aware regularization, which enforces stronger early-stage constraints to preserve diversity without compromising reward optimization efficiency. Experiments demonstrate that our method achieves a 13\%--18\% improvement in semantic diversity under matched quality scores, establishing a new Pareto frontier between image quality and diversity for GRPO-based image generation.

图像生成强化学习多样性GRPO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。