用多智能体协作解析复杂提示,生成更精准的图像。
MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image Generation
- 设计多智能体系统,分工解析场景元素。
- 分层组合扩散提升区域细节,生成高保真图像。
- 无需训练即可显著提升复杂场景生成效果。
扩散模型在文本到图像生成中表现优异,但在处理包含多个物体、属性和关系的复杂提示时仍存在性能瓶颈。为此,我们提出基于多智能体协作的组合扩散方法(MCCD),用于复杂场景生成。具体而言,设计了一种多智能体协作式场景解析模块,利用多模态大语言模型(MLLMs)有效提取各类场景元素,构建具有不同任务的智能体系统。同时,采用分层组合扩散机制,通过高斯掩码与过滤策略细化边界框区域,并增强对象特征,从而实现复杂场景的精准、高质量生成。大量实验表明,MCCD以无训练方式显著提升基线模型性能,在复杂场景生成上具有明显优势。
原文摘要 · Abstract (English)
Diffusion models have shown excellent performance in text-to-image generation. Nevertheless, existing methods often suffer from performance bottlenecks when handling complex prompts that involve multiple objects, characteristics, and relations. Therefore, we propose a Multi-agent Collaboration-based Compositional Diffusion (MCCD) for text-to-image generation for complex scenes. Specifically, we design a multi-agent collaboration-based scene parsing module that generates an agent system comprising multiple agents with distinct tasks, utilizing MLLMs to extract various scene elements effectively. In addition, Hierarchical Compositional diffusion utilizes a Gaussian mask and filtering to refine bounding box regions and enhance objects through region enhancement, resulting in the accurate and high-fidelity generation of complex scenes. Comprehensive experiments demonstrate that our MCCD significantly improves the performance of the baseline models in a training-free manner, providing a substantial advantage in complex scene generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。