让多个物体在生成中位置精准、身份一致,解决多实例图像生成难题
ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation
- 用布局图和参考图共同引导,通过上下文锚定确保物体位置准确
- 在复杂布局下保持多主体身份一致性,生成效果优于现有方法
- 适合需要高精度控制的图像生成任务,如广告设计与虚拟场景构建
多实例图像生成(MIG)对现代扩散模型仍是重大挑战,主要受限于对物体布局的精确控制能力不足以及多个不同主体的身份保持困难。为此,我们提出 ContextGen,一种基于布局和参考图像引导的扩散变压器框架。该方法引入两项关键技术:上下文布局锚定(CLA)机制,将组合布局图融入生成上下文,以鲁棒地锚定物体在期望位置;身份一致性注意力(ICA),利用上下文参考图确保多个实例的身份一致性。为解决该任务缺乏大规模高质量数据集的问题,我们构建了 IMIG-100K,首个专为多实例生成设计、包含详细布局与身份标注的数据集。大量实验表明,ContextGen 达到新基准,尤其在布局控制与身份保真度方面显著优于现有方法。
原文摘要 · Abstract (English)
Multi-instance image generation (MIG) remains a significant challenge for modern diffusion models due to key limitations in achieving precise control over object layout and preserving the identity of multiple distinct subjects. To address these limitations, we introduce ContextGen, a novel Diffusion Transformer framework for multi-instance generation that is guided by both layout and reference images. Our approach integrates two key technical contributions: a Contextual Layout Anchoring (CLA) mechanism that incorporates the composite layout image into the generation context to robustly anchor the objects in their desired positions, and Identity Consistency Attention (ICA), an innovative attention mechanism that leverages contextual reference images to ensure the identity consistency of multiple instances. To address the absence of a large-scale, high-quality dataset for this task, we introduce IMIG-100K, the first dataset to provide detailed layout and identity annotations specifically designed for Multi-Instance Generation. Extensive experiments demonstrate that ContextGen sets a new state-of-the-art, outperforming existing methods especially in layout control and identity fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。