arXiv:2505.24086cs.CV2025-05被引 1

用大模型生成带深度信息的布局,提升复杂物体组合的图像生成质量。

ComposeAnything: Composite Object Priors for Text-to-Image Generation

  • 用大模型生成含深度信息的2.5D语义布局,替代随机噪声初始化。
  • 在T2I-CompBench和NSR-1K上优于现有方法,支持高数量与超现实组合。
  • 适合需要精确物体位置和复杂构图的图像生成任务。

当前文本到图像(T2I)模型在生成涉及复杂且新颖物体排列的图像时仍面临挑战。尽管基于布局的方法通过二维布局的空间约束提升了物体排布,但往往难以捕捉三维位置信息,且牺牲了图像质量和连贯性。本文提出ComposeAnything,一种无需重训练现有T2I模型即可提升组合图像生成能力的新框架。该方法首先利用大语言模型(LLM)的思维链能力,从文本生成包含深度信息的2.5D语义布局,即带有深度信息的2D物体边界框与详细描述。基于此布局,生成空间与深度感知的粗略复合图像,作为扩散模型中强而可解释的先验,替代随机噪声初始化。该先验通过对象先验强化与空间控制去噪引导去噪过程,实现组合物体与连贯背景的无缝生成,并支持对不准确先验的迭代优化。在包含2D/3D空间排列、高物体数量及超现实组合的T2I-CompBench与NSR-1K基准测试中,本方法均优于当前最优技术。人类评估进一步验证其生成图像质量高,且构图忠实于文本描述。

原文摘要 · Abstract (English)

Generating images from text involving complex and novel object arrangements remains a significant challenge for current text-to-image (T2I) models. Although prior layout-based methods improve object arrangements using spatial constraints with 2D layouts, they often struggle to capture 3D positioning and sacrifice quality and coherence. In this work, we introduce ComposeAnything, a novel framework for improving compositional image generation without retraining existing T2I models. Our approach first leverages the chain-of-thought reasoning abilities of LLMs to produce 2.5D semantic layouts from text, consisting of 2D object bounding boxes enriched with depth information and detailed captions. Based on this layout, we generate a spatial and depth aware coarse composite of objects that captures the intended composition, serving as a strong and interpretable prior that replaces stochastic noise initialization in diffusion-based T2I models. This prior guides the denoising process through object prior reinforcement and spatial-controlled denoising, enabling seamless generation of compositional objects and coherent backgrounds, while allowing refinement of inaccurate priors. ComposeAnything outperforms state-of-the-art methods on the T2I-CompBench and NSR-1K benchmarks for prompts with 2D/3D spatial arrangements, high object counts, and surreal compositions. Human evaluations further demonstrate that our model generates high-quality images with compositions that faithfully reflect the text.

文本生成图像组合生成2.5D布局扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。