arXiv:2509.25562cs.AIcs.CL2025-09被引 1

用内在奖励提升图像生成质量,无需人工标注数据。

IRIS: Intrinsic Reward Image Synthesis

  • 通过降低模型自信心来优化生成效果,反直觉但有效。
  • 在无外部奖励下,生成图像细节更丰富,符合人类偏好。
  • 适合追求高质量图像生成且缺乏标注数据的研究者。

尽管强化学习从人类反馈(RLHF)在语言推理中取得成功,但在自回归文本到图像(T2I)生成中的应用常受限于人类偏好数据的稀缺。本文探讨了自回归T2T模型如何仅依靠内部信号学习,而无需外部奖励或标注数据。与数学和代码推理中最大化自信心的发现相反,我们发现降低自信心反而能提升图像生成质量。观察表明,高自信心模型倾向于生成简单、单一的图像,与人类偏好不符;而低自信心模型则更可能生成细节丰富、生动的图像。基于此,我们提出IRIS(Intrinsic Reward Image Synthesis),首个仅使用内在奖励改进自回归T2I模型的框架。实验表明,将IRIS应用于自回归T2I模型,其性能优于单独使用外部奖励的模型,甚至媲美集成多个外部奖励的模型。IRIS还促进了高质量图像生成所需的精细思维链(CoT)推理的涌现。

原文摘要 · Abstract (English)

Despite the success of Reinforcement Learning from Human Feedback (RLHF) in language reasoning, its application to autoregressive Text-to-Image (T2I) generation is often constrained by the limited availability of human preference data. This paper explores how an autoregressive T2I model can learn from internal signals without relying on external rewards or labeled data. Contrary to recent findings in math and code reasoning, we show that minimizing self-certainty, rather than maximizing it, improves image generation. We observe that autoregressive T2I models with higher certainty are likely to generate simple and uniform images, which are less aligned with human preferences, and models with lower certainty are likely to generate vivid images rich in detail. Based on this observation, we propose IRIS(Intrinsic Reward Image Synthesis), the first framework to improve autoregressive T2I models with reinforcement learning using only an intrinsic reward. Empirical results demonstrate that applying IRIS to autoregressive T2I models achieves performance superior to those trained by individual external rewards, and matching those trained by ensemble external rewards. IRIS also incentivizes the emergence of nuanced CoT reasoning for high-quality image generation.

图像生成强化学习内在奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。