用草图引导视觉思维,提升文本生图精度与稀有概念生成能力
DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
- 先生成低分辨率草图作为视觉规划,再通过模型验证并修正语义偏差
- 在GenEval等基准上提升8%~9.1%,显著优于传统文本推理生成方法
- 适合需要高精度控制或罕见元素组合的图像生成任务
近期统一多模态大模型虽具备链式思考(CoT)能力,但在文本到图像生成中仍受限于抽象文本规划或单一生成模式。为此,我们提出Draft-as-CoT(DraCo),一种融合文本与视觉内容的交错推理范式。首先生成低分辨率草图作为预览,提供更具体、结构化的视觉指导;随后利用模型理解能力检测草图与输入提示间的语义偏差,并通过选择性修正与超分辨率实现优化。该方法解决文本规划粗糙和罕见属性组合难生成两大难题。为支持训练,我们构建了包含24万样本的DraCo-240K数据集,涵盖通用修正、实例操作与布局重组三类基础能力。结合专为交错推理设计的DraCo-CFG无分类器引导策略,DraCo在GenEval(+8%)、Imagine-Bench(+0.91)和GenEval++(+3%)上表现优异,显著超越直接生成及其它基于CoT的方法。
原文摘要 · Abstract (English)
Recent unified multimodal large language models (MLLMs) have shown impressive capabilities, incorporating chain-of-thought (CoT) reasoning for enhanced text-to-image generation. However, existing approaches remain limited, either treating the model merely as a standalone generator or relying on abstract textual planning. To this end, we propose Draft-as-CoT (DraCo), a novel interleaved reasoning paradigm that fully leverages both textual and visual contents in CoT for better planning and verification. Our method first generates a low-resolution draft image as preview, providing more concrete and structural visual planning and guidance. Then, we employ the model's inherent understanding capability to verify potential semantic misalignments between the draft and input prompt, and performs refinement through selective corrections with super-resolution. In this way, our approach addresses two fundamental challenges: the coarse-grained nature of textual planning and the difficulty in generating rare attribute combinations. To support training, we curate DraCo-240K, aiming to enhance three atomic capabilities spanning general correction, instance manipulation, and layout reorganization. Supported by DraCo-CFG, a specialized classifier-free guidance (CFG) strategy for interleaved reasoning, DraCo achieves a tremendous increase on GenEval (+8%), Imagine-Bench (+0.91), and GenEval++ (+3%), significantly outperforming direct generation and other generation methods empowered by CoT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。