arXiv:2501.09194cs.CVcs.AI2025-01被引 3

让图像生成精准按指定位置和对象布局,提升可控性与质量。

Grounding Text-to-Image Diffusion Models for Controlled High-Quality Image Generation

  • 结合控制网与文本定位技术,用框定位置生成物体。
  • 在COCO数据集上实现46.6的精度、44.5的召回率和19.8的FID。
  • 适合需要高精度布局控制的图像生成场景。

文本到图像生成扩散模型在根据文本描述生成多样化高质量图像方面表现出色。已有布局到图像模型利用分割图、边缘、人体关键点等布局信息来控制生成过程。本文提出ObjectDiffusion,通过引入语义与空间定位信息,使扩散模型能够精确渲染并放置目标物体于指定边界框位置。为此,我们在ControlNet架构基础上进行大幅修改,并融合GLIGEN提出的定位方法。模型在COCO2017训练集上微调,并在验证集上评估。结果表明,该模型在可控图像生成的精度与质量上均优于当前基于开源数据集的SOTA模型,达到AP₅₀为46.6,AR为44.5,FID为19.8。ObjectDiffusion展现出在多种语境下对封闭集与开放集词汇的强定位能力,可生成多类、多尺寸、多形态且位置准确的细节丰富图像。

原文摘要 · Abstract (English)

Text-to-image (T2I) generative diffusion models have demonstrated outstanding performance in synthesizing diverse, high-quality visuals from text captions. Several layout-to-image models have been developed to control the generation process by utilizing a wide range of layouts, such as segmentation maps, edges, and human keypoints. In this work, we propose ObjectDiffusion, a model that conditions T2I diffusion models on semantic and spatial grounding information, enabling the precise rendering and placement of desired objects in specific locations defined by bounding boxes. To achieve this, we make substantial modifications to the network architecture introduced in ControlNet to integrate it with the grounding method proposed in GLIGEN. We fine-tune ObjectDiffusion on the COCO2017 training dataset and evaluate it on the COCO2017 validation dataset. Our model improves the precision and quality of controllable image generation, achieving an AP$_{\text{50}}$ of 46.6, an AR of 44.5, and an FID of 19.8, outperforming the current SOTA model trained on open-source datasets across all three metrics. ObjectDiffusion demonstrates a distinctive capability in synthesizing diverse, high-quality, high-fidelity images that seamlessly conform to the semantic and spatial control layout. Evaluated in qualitative and quantitative tests, ObjectDiffusion exhibits remarkable grounding capabilities in closed-set and open-set vocabulary settings across a wide variety of contexts. The qualitative assessment verifies the ability of ObjectDiffusion to generate multiple detailed objects in varying sizes, forms, and locations.

图像生成扩散模型可控生成空间定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。