arXiv:2511.18378cs.CV2025-11被引 1

用合成课程强化学习提升文本生成图像的组合能力

Synthetic Curriculum Reinforces Compositional Text-to-Image Generation

  • 基于场景图设计难度评估,自适应生成渐进式训练数据
  • 在扩散模型和自回归模型上显著提升复杂场景生成效果
  • 适合研究文本到图像生成与强化学习融合的研究者

文本到图像(T2I)生成长期面临挑战,尤其在组合性合成方面。该任务需准确渲染包含多个对象、多样属性及复杂空间与语义关系的场景,要求精确的对象布局和连贯的交互。本文提出一种名为CompGen的新颖组合式课程强化学习框架,利用场景图建立组合能力的难度评估标准,并开发相应的自适应马尔可夫链蒙特卡洛图采样算法。该难度感知方法可生成逐步优化的训练课程数据,通过强化学习持续提升T2I模型性能。我们将课程学习整合至组相对策略优化(GRPO),并探索不同课程调度策略。实验表明,CompGen在不同调度策略下呈现明显缩放曲线,易到难及高斯采样策略优于随机采样。大量实验证明,CompGen显著增强基于扩散模型与自回归模型的组合生成能力,验证了其在提升组合式T2I系统中的有效性。

原文摘要 · Abstract (English)

Text-to-Image (T2I) generation has long been an open problem, with compositional synthesis remaining particularly challenging. This task requires accurate rendering of complex scenes containing multiple objects that exhibit diverse attributes as well as intricate spatial and semantic relationships, demanding both precise object placement and coherent inter-object interactions. In this paper, we propose a novel compositional curriculum reinforcement learning framework named CompGen that addresses compositional weakness in existing T2I models. Specifically, we leverage scene graphs to establish a novel difficulty criterion for compositional ability and develop a corresponding adaptive Markov Chain Monte Carlo graph sampling algorithm. This difficulty-aware approach enables the synthesis of training curriculum data that progressively optimize T2I models through reinforcement learning. We integrate our curriculum learning approach into Group Relative Policy Optimization (GRPO) and investigate different curriculum scheduling strategies. Our experiments reveal that CompGen exhibits distinct scaling curves under different curriculum scheduling strategies, with easy-to-hard and Gaussian sampling strategies yielding superior scaling performance compared to random sampling. Extensive experiments demonstrate that CompGen significantly enhances compositional generation capabilities for both diffusion-based and auto-regressive T2I models, highlighting its effectiveness in improving the compositional T2I generation systems.

文本生成图像组合生成强化学习课程学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。