让AI画画前先思考构图和操作,更懂人类意图。
GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing
- 用语言推理链规划图像生成与编辑的语义和空间关系
- 构建900万条带推理链的数据集,显著提升生成质量
- 支持交互式修改推理步骤,适合需要精准控制的创作者
当前图像生成与编辑方法主要将文本提示直接输入,缺乏对视觉构图和显式操作的推理。我们提出生成思维链(GoT),一种在输出图像前通过显式语言推理过程实现生成与编辑的新范式。该方法将传统文本到图像生成转变为推理引导框架,分析语义关系与空间布局。我们定义了GoT的形式化表达,并构建包含超过900万样本的大规模GoT数据集,其中包含详细推理链以捕捉语义-空间关系。为充分利用GoT优势,我们实现统一框架,结合Qwen2.5-VL生成推理链,并采用新型语义-空间引导模块增强端到端扩散模型。实验表明,我们的GoT框架在生成与编辑任务上表现优异,相比基线有显著提升。此外,该方法支持交互式视觉生成,允许用户显式修改推理步骤以精确调整图像。GoT开创了推理驱动视觉生成与编辑的新方向,生成结果更契合人类意图。为促进后续研究,我们已将数据集、代码及预训练模型公开于https://github.com/rongyaofang/GoT。
原文摘要 · Abstract (English)
Current image generation and editing methods primarily process textual prompts as direct inputs without reasoning about visual composition and explicit operations. We present Generation Chain-of-Thought (GoT), a novel paradigm that enables generation and editing through an explicit language reasoning process before outputting images. This approach transforms conventional text-to-image generation and editing into a reasoning-guided framework that analyzes semantic relationships and spatial arrangements. We define the formulation of GoT and construct large-scale GoT datasets containing over 9M samples with detailed reasoning chains capturing semantic-spatial relationships. To leverage the advantages of GoT, we implement a unified framework that integrates Qwen2.5-VL for reasoning chain generation with an end-to-end diffusion model enhanced by our novel Semantic-Spatial Guidance Module. Experiments show our GoT framework achieves excellent performance on both generation and editing tasks, with significant improvements over baselines. Additionally, our approach enables interactive visual generation, allowing users to explicitly modify reasoning steps for precise image adjustments. GoT pioneers a new direction for reasoning-driven visual generation and editing, producing images that better align with human intent. To facilitate future research, we make our datasets, code, and pretrained models publicly available at https://github.com/rongyaofang/GoT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。