arXiv:2503.01298cs.CVcs.AI2025-03被引 8

让生成模型像人一样分步思考,提升复杂图像生成能力

Towards Enhanced Image Generation Via Multi-modal Chain of Thought in Unified Generative Models

  • 设计功能导向的专家并行结构,分离模态冲突
  • 提出多模态思维链,模拟规划-执行-反思-修正流程
  • 无需完整多步数据,通过联合训练实现端到端优化

统一生成模型在文本与图像生成中表现卓越。图像合成任务通常采用直接的文本到图像(T2I)生成方式,但这种方式难以处理现实中常见的复杂组合指令。尽管该问题至关重要,现有工作主要聚焦于提升基础生成能力,仍无法有效解决复杂指令生成难题。受思维链(CoT)逐步解题启发,本文将CoT引入统一生成模型,以增强其应对复杂图像生成的能力。为此,提出功能导向专家(FoXperts),一种按功能分配专家的并行架构,摆脱主流模态导向设计中的潜在冲突,为思维链提供基础支持。针对多模态思维链(MCoT)的设计挑战,我们模拟人类艺术创作流程——规划、执行、反思与修正,构建了适用于图文双重输入的MCoT方法。为克服多步数据难以收集的问题,设计了一种多任务联合训练方案,以解耦方式赋予模型各步骤所需能力。大量实验表明,FoX在多个T2I基准上持续优于现有统一模型,在复杂图像生成方面取得显著提升。

原文摘要 · Abstract (English)

Unified generative models have shown remarkable performance in text and image generation. For image synthesis tasks, they adopt straightforward text-to-image (T2I) generation. However, direct T2I generation limits the models in handling complex compositional instructions, which frequently occur in real-world scenarios. Although this issue is vital, existing works mainly focus on improving the basic image generation capability of the models. While such improvements help to some extent, they still fail to adequately resolve the problem. Inspired by Chain of Thought (CoT) solving complex problems step by step, this work aims to introduce CoT into unified generative models to address the challenges of complex image generation that direct T2I generation cannot effectively solve, thereby endowing models with enhanced image generation ability. To achieve this, we first propose Functionality-oriented eXperts (FoXperts), an expert-parallel architecture in our model FoX, which assigns experts by function. FoXperts disentangles potential conflicts in mainstream modality-oriented designs and provides a solid foundation for CoT. When introducing CoT, the first question is how to design it for complex image generation. To this end, we emulate a human-like artistic workflow--planning, acting, reflection, and correction--and propose the Multimodal Chain of Thought (MCoT) approach, as the data involves both text and image. To address the subsequent challenge of designing an effective MCoT training paradigm, we develop a multi-task joint training scheme that equips the model with all capabilities required for each MCoT step in a disentangled manner. This paradigm avoids the difficulty of collecting consistent multi-step data tuples. Extensive experiments show that FoX consistently outperforms existing unified models on various T2I benchmarks, delivering notable improvements in complex image generation.

图像生成多模态思维链统一模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。