用图像生成结构+语言指令,精准控制多个物体的属性和位置。
InstanceGen: Image Generation with Instance-level Instructions
- 用生成图像提供细粒度结构初始化,结合大模型指令
- 能准确还原对象数量、实例属性和空间关系
- 适合需要精细控制生成内容的设计师或研究者
尽管生成模型能力快速提升,预训练文本到图像模型在处理包含多个对象和实例级属性的复杂提示时仍存在语义理解困难。为此,研究者开始引入粗略边界框等结构约束来增强生成控制。本文进一步推进该思路,发现现代图像生成模型可直接提供合理的细粒度结构初始化。我们提出一种方法,将基于图像的结构引导与基于大语言模型的实例级指令相结合,使生成图像严格遵循文本提示中的所有要求,包括对象数量、实例属性以及实例间的空间关系。
原文摘要 · Abstract (English)
Despite rapid advancements in the capabilities of generative models, pretrained text-to-image models still struggle in capturing the semantics conveyed by complex prompts that compound multiple objects and instance-level attributes. Consequently, we are witnessing growing interests in integrating additional structural constraints, typically in the form of coarse bounding boxes, to better guide the generation process in such challenging cases. In this work, we take the idea of structural guidance a step further by making the observation that contemporary image generation models can directly provide a plausible fine-grained structural initialization. We propose a technique that couples this image-based structural guidance with LLM-based instance-level instructions, yielding output images that adhere to all parts of the text prompt, including object counts, instance-level attributes, and spatial relations between instances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。