arXiv:2509.06945cs.CVcs.AI2025-09被引 37

通过交替思考与生成,提升文本到图像的细节还原与指令遵循能力。

Interleaving Reasoning for Better Text-to-Image Generation

  • 采用文本思考与图像生成交替的框架,逐步优化图像细节。
  • 在多个评测中取得领先,关键指标提升5-10分。
  • 适合关注图像生成质量与语义准确性的研究者使用。

统一的多模态理解与生成模型在图像生成方面已取得显著进展,但在指令遵循和细节保留方面仍明显落后于如GPT-4o等紧耦合理解与生成的系统。受最近交错推理进展的启发,我们探索其在文本到图像(T2I)生成中的应用。提出交错推理生成(IRG)框架,交替进行文本思考与图像合成:模型先生成文本思考以指导初始图像,再基于结果反思,优化细粒度细节、视觉质量和美学表现,同时保持语义一致。为有效训练IRG,提出交错推理生成学习(IRGL),聚焦两个子目标:(1)强化初始思考-生成阶段,建立核心内容与基础质量;(2)实现高质量文本反思,并在后续图像中忠实执行修正。构建了包含六种分解学习模式的IRGL-300K数据集,覆盖文本思考与完整思考-图像轨迹。从原生支持交错输出的基础模型出发,采用两阶段训练:先建立稳健的思考与反思能力,再高效微调全轨迹数据上的IRG流程。大量实验表明,该方法在GenEval、WISE、TIIF、GenAI-Bench和OneIG-EN上均达到最先进水平,绝对提升5-10分,同时显著改善视觉质量与细粒度保真度。代码、模型权重与数据集将公开于:https://github.com/Osilly/Interleaving-Reasoning-Generation。

原文摘要 · Abstract (English)

Unified multimodal understanding and generation models recently have achieve significant improvement in image generation capability, yet a large gap remains in instruction following and detail preservation compared to systems that tightly couple comprehension with generation such as GPT-4o. Motivated by recent advances in interleaving reasoning, we explore whether such reasoning can further improve Text-to-Image (T2I) generation. We introduce Interleaving Reasoning Generation (IRG), a framework that alternates between text-based thinking and image synthesis: the model first produces a text-based thinking to guide an initial image, then reflects on the result to refine fine-grained details, visual quality, and aesthetics while preserving semantics. To train IRG effectively, we propose Interleaving Reasoning Generation Learning (IRGL), which targets two sub-goals: (1) strengthening the initial think-and-generate stage to establish core content and base quality, and (2) enabling high-quality textual reflection and faithful implementation of those refinements in a subsequent image. We curate IRGL-300K, a dataset organized into six decomposed learning modes that jointly cover learning text-based thinking, and full thinking-image trajectories. Starting from a unified foundation model that natively emits interleaved text-image outputs, our two-stage training first builds robust thinking and reflection, then efficiently tunes the IRG pipeline in the full thinking-image trajectory data. Extensive experiments show SoTA performance, yielding absolute gains of 5-10 points on GenEval, WISE, TIIF, GenAI-Bench, and OneIG-EN, alongside substantial improvements in visual quality and fine-grained fidelity. The code, model weights and datasets will be released in: https://github.com/Osilly/Interleaving-Reasoning-Generation .

文本生成图像推理机制图像质量多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。