用场景图结构提升复杂文本生成图像的准确性与多样性
All-in-One Conditioning for Text-to-Image Synthesis
- 基于场景图构建零样本视觉引导机制
- 通过轻量语言模型生成属性-尺寸-数量-位置条件
- 适合需要高语义对齐和灵活构图的图像生成任务
复杂提示中包含多个对象、属性和空间关系的准确理解与可视化,是文本到图像合成中的关键挑战。尽管现有模型能生成逼真图像,但在处理复杂文本输入时仍难以保持语义一致性和结构连贯性。本文提出一种基于场景图的新方法,将文本到图像合成置于场景图框架内,以增强模型的组合能力。不同于依赖预定义布局图的刚性约束,本方法引入零样本、基于场景图的条件机制,在推理时生成软性视觉引导。核心为轻量级的属性-尺寸-数量-位置(ASQL)条件器,通过轻量语言模型生成视觉条件,并在推理时优化扩散模型生成过程。该机制在保持文本-图像对齐的同时,支持轻量、一致且多样化的图像生成。
原文摘要 · Abstract (English)
Accurate interpretation and visual representation of complex prompts involving multiple objects, attributes, and spatial relationships is a critical challenge in text-to-image synthesis. Despite recent advancements in generating photorealistic outputs, current models often struggle with maintaining semantic fidelity and structural coherence when processing intricate textual inputs. We propose a novel approach that grounds text-to-image synthesis within the framework of scene graph structures, aiming to enhance the compositional abilities of existing models. Eventhough, prior approaches have attempted to address this by using pre-defined layout maps derived from prompts, such rigid constraints often limit compositional flexibility and diversity. In contrast, we introduce a zero-shot, scene graph-based conditioning mechanism that generates soft visual guidance during inference. At the core of our method is the Attribute-Size-Quantity-Location (ASQL) Conditioner, which produces visual conditions via a lightweight language model and guides diffusion-based generation through inference-time optimization. This enables the model to maintain text-image alignment while supporting lightweight, coherent, and diverse image synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。