让文字直接控制图像布局,实现精准多对象生成
ConsistCompose: Unified Multimodal Layout Control for Image Composition
- 将坐标嵌入文本提示,用语言直接指挥物体位置
- 在两个数据集上空间准确率显著优于基线模型
- 适合需要精确布局控制的图像生成任务
统一多模态模型虽快速进展,但多数仍聚焦视觉定位,而语言嵌入布局控制(LELG)的多实例生成仍待探索。我们提出 ConsistCompose 框架,将布局坐标直接嵌入语言提示中,通过单一生成接口实现基于交错图文输入的可控多实例图像生成。构建了包含340万条样本的 ConsistCompose3M 数据集(260万文本引导,80万图像引导),提供大规模布局监督。利用实例-坐标绑定提示与坐标感知无分类器引导,将语言布局线索转化为精确空间控制,无需特定任务分支。在 COCO-Position 与 MS-Bench 上实验表明,ConsistCompose 显著提升空间准确性,同时保持身份一致性与竞争力的通用多模态理解能力,建立统一的布局可控多模态图像生成范式。
原文摘要 · Abstract (English)
Unified multimodal models that couple visual understanding with image generation have advanced rapidly, yet most systems still focus on visual grounding-aligning language with image regions-while their generative counterpart, linguistic-embedded layout-grounded generation (LELG) for layout-controllable multi-instance generation, remains underexplored and limits precise compositional control. We present ConsistCompose, a unified multimodal framework that embeds layout coordinates directly into language prompts, enabling layout-controlled multi-instance image generation from Interleaved Image-Text within a single generative interface. We further construct ConsistCompose3M, a 3.4M multi-instance generation dataset with layout and identity annotations (2.6M text-guided and 0.8M image-guided data pairs) that provides large-scale supervision for layout-conditioned generation. Within this framework, LELG is instantiated through instance-coordinate binding prompts and coordinate-aware classifier-free guidance, which translate linguistic layout cues into precise spatial control without task-specific branches. Experiments on COCO-Position and MS-Bench show that ConsistCompose substantially improves spatial accuracy over layout-controlled baselines while preserving identity fidelity and competitive general multimodal understanding, establishing a unified paradigm for layout-controllable multimodal image generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。