arXiv:2505.17540cs.CVcs.AI2025-05被引 17

用强化学习让AI自己优化提示词,生成更符合意图的图像。

RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning

  • 通过强化学习训练语言模型生成带自我反思的结构化提示词
  • 在GenEval和T2I-Compbench上提升空间布局准确性和组合泛化能力
  • 无需人工标注数据,适合需要精准控制生成内容的研究者

尽管文本到图像生成技术取得进展,现有模型仍难以忠实捕捉简短、不明确提示中的用户意图。以往方法虽尝试利用大语言模型增强提示词,但常因缺乏视觉语义和真实构图的约束,生成风格化或不现实的内容。受语言模型推理能力启发,我们提出RePrompt,一种通过强化学习将显式推理引入提示词优化的新框架。该方法训练语言模型生成结构化、自我反思的提示词,以图像级结果为优化目标。定制的奖励模型从人类偏好、语义对齐和视觉构图三方面评估生成图像,提供间接监督以改进提示生成。本方法实现端到端训练,无需人工标注数据。在GenEval和T2I-Compbench上的实验表明,RePrompt显著提升了空间布局保真度和组合泛化能力,适用于多种主流文本到图像生成模型,并达到新基准水平。

原文摘要 · Abstract (English)

Despite recent progress in text-to-image (T2I) generation, existing models often struggle to faithfully capture user intentions from short and under-specified prompts. While prior work has attempted to enhance prompts using large language models (LLMs), these methods frequently generate stylistic or unrealistic content due to insufficient grounding in visual semantics and real-world composition. Inspired by recent advances in reasoning for language model, we propose RePrompt, a novel reprompting framework that introduces explicit reasoning into the prompt enhancement process via reinforcement learning. Instead of relying on handcrafted rules or stylistic rewrites, our method trains a language model to generate structured, self-reflective prompts by optimizing for image-level outcomes. The tailored reward models assesse the generated images in terms of human preference, semantic alignment, and visual composition, providing indirect supervision to refine prompt generation. Our approach enables end-to-end training without human-annotated data. Experiments on GenEval and T2I-Compbench show that RePrompt significantly boosts spatial layout fidelity and compositional generalization across diverse T2I backbones, establishing new state-of-the-art results.

文本生成图像强化学习提示工程视觉对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。