通过分阶段奖励提升文本生成图像的多属性理解能力
Visual-CoG: Stage-Aware Reinforcement Learning with Chain of Guidance for Text-to-Image Generation
- 引入三阶段视觉推理框架,每阶段提供即时奖励
- 在多个基准上分别提升15%、5%、19%性能
- 适合需要精准控制生成过程的研究者
尽管近期自回归模型在文本到图像(T2I)生成方面取得进展,其对多属性和模糊提示的处理能力仍有限。现有方法虽采用思维链(CoT)实现分阶段视觉合成,并用强化学习增强推理能力,但多数仅在生成末尾提供奖励信号。这种单一终局反馈难以定位各阶段贡献,易导致策略次优。为此,我们提出视觉思维链(Visual-CoG)范式,包含语义推理、过程精炼与结果评估三个阶段,全程提供分阶段奖励以实现即时引导。我们还构建了视觉认知基准VisCog-Bench,包含四项子任务用于评估语义推理效果。在GenEval、T2I-CompBench及所提VisCog-Bench上的综合评估显示,性能分别提升15%、5%和19%,验证了Visual-CoG的优越性。相关资源将很快公开。
原文摘要 · Abstract (English)
Despite the promising progress of recent autoregressive models in text-to-image (T2I) generation, their ability to handle multi-attribute and ambiguous prompts remains limited. To address these limitations, existing works have applied chain-of-thought (CoT) to enable stage-aware visual synthesis and employed reinforcement learning (RL) to improve reasoning capabilities. However, most models provide reward signals only at the end of the generation stage. This monolithic final-only guidance makes it difficult to identify which stages contribute positively to the final outcome and may lead to suboptimal policies. To tackle this issue, we propose a Visual-Chain of Guidance (Visual-CoG) paradigm consisting of three stages: semantic reasoning, process refining, and outcome evaluation, with stage-aware rewards providing immediate guidance throughout the image generation pipeline. We further construct a visual cognition benchmark, VisCog-Bench, which comprises four subtasks to evaluate the effectiveness of semantic reasoning. Comprehensive evaluations on GenEval, T2I-CompBench, and the proposed VisCog-Bench show improvements of 15%, 5%, and 19%, respectively, demonstrating the superior performance of the proposed Visual-CoG. We will release all the resources soon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。