用多模态大模型实现文本到图像生成的精准空间控制
LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs
- 用多模态大模型做全局布局规划和语义增强
- 在多个T2I模型上实现优于现有方法的生成质量
- 适合需要复杂场景控制的图像生成任务
当前基于扩散模型的文本到图像生成方法在空间控制方面已有显著进展,但仍难以精确遵循复杂的文本指令,尤其在涉及多个对象或复杂空间关系时。本文提出一种名为LLMControl的框架,利用多模态大模型作为全局控制器,提升对齐能力,通过融合视觉条件与文本提示,互补地调控结构与外观生成。该框架能自动安排空间布局、增强语义描述并绑定对象属性,生成的控制信号被注入去噪网络,以根据新采样约束重新聚焦和强化注意力图。大量定性与定量实验表明,LLMControl在多个预训练T2I模型上均达到领先性能,尤其在处理复杂输入条件时表现突出,多数现有方法在此类场景下失败。
原文摘要 · Abstract (English)
Recent spatial control methods for text-to-image (T2I) diffusion models have shown compelling results. However, these methods still fail to precisely follow the control conditions and generate the corresponding images, especially when encountering the textual prompts that contain multiple objects or have complex spatial compositions. In this work, we present a LLM-guided framework called LLM\_Control to address the challenges of the controllable T2I generation task. By improving grounding capabilities, LLM\_Control is introduced to accurately modulate the pre-trained diffusion models, where visual conditions and textual prompts influence the structures and appearance generation in a complementary way. We utilize the multimodal LLM as a global controller to arrange spatial layouts, augment semantic descriptions and bind object attributes. The obtained control signals are injected into the denoising network to refocus and enhance attention maps according to novel sampling constraints. Extensive qualitative and quantitative experiments have demonstrated that LLM\_Control achieves competitive synthesis quality compared to other state-of-the-art methods across various pre-trained T2I models. It is noteworthy that LLM\_Control allows the challenging input conditions on which most of the existing methods
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。