arXiv:2604.02355cs.LGcs.CV2026-04被引 9

通过熵调控优化生成过程,提升自回归图像生成的稳定性和质量。

From Broad Exploration to Stable Synthesis: Entropy-Guided Optimization for Autoregressive Image Generation

论文配图:From Broad Exploration to Stable Synthesis: Entropy-Guided Optimization for Autoregressive Image Generation
图 1 · 摘自论文原文
  • 基于熵分析揭示思维链与强化学习的交互机制
  • 降低图像标记熵可显著提升生成质量,且与奖励强负相关
  • 针对高不确定性标记动态调整优化策略,防止崩溃

将思维链(CoT)与强化学习(RL)结合可提升文本到图像(T2I)生成效果,但其内在交互机制尚不明确。本文通过系统的熵基分析,发现:(1) CoT扩展生成探索空间,而RL将其收缩至高奖励区域;(2) 最终奖励与图像标记熵的均值和方差呈强负相关,表明需减少不确定性和不稳定性;(3) 文本思维链的熵直接决定下游图像质量,低熵思维链生成效果更优。受此启发,提出熵引导的组相对策略优化(EG-GRPO),通过不确定性重分配优化预算:对低熵标记排除奖励驱动更新以保持稳定,对高熵标记给予熵奖励以鼓励结构化探索而不致坍塌。在标准T2I基准上实验表明,EG-GRPO达到当前最优性能。

原文摘要 · Abstract (English)

Combining Chain-of-Thought (CoT) with Reinforcement Learning (RL) improves text-to-image (T2I) generation, yet the underlying interaction between CoT's exploration and RL's optimization remains unclear. We present a systematic entropy-based analysis that yields three key insights: (1) CoT expands the generative exploration space, while RL contracts it toward high-reward regions; (2) final reward is strongly negatively correlated with both the mean and variance of image-token entropy, highlighting the need to reduce uncertainty and instability; and (3) the entropy of the textual CoT directly governs downstream image quality, with lower-entropy CoTs leading to better generations. Motivated by these findings, we propose Entropy-Guided Group Relative Policy Optimization (EG-GRPO), a fine-tuning strategy that reallocates optimization budget by uncertainty: low-entropy tokens are excluded from reward-driven updates to preserve stability, while high-entropy tokens receive an entropy bonus that encourages structured exploration without collapse. Experiments on standard T2I benchmarks demonstrate that EG-GRPO achieves state-of-the-art performance.

图像生成强化学习熵控制自回归模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。