无需微调即可精准控制图像中多个对象的位置和外观。
LocRef-Diffusion:Tuning-Free Layout and Appearance-Guided Generation
- 用布局网络结合显式布局与跨注意力,精准定位对象位置。
- 通过外观网络提取参考图特征并融入扩散模型,提升外观还原度。
- 适用于需要个性定制图像生成的设计师或内容创作者。
基于扩散的文生图模型近年来在生成高质量图像方面取得了显著进展。然而,如何实现图像中多个实例的个性化、可控制生成仍是亟待解决的问题。本文提出LocRef-Diffusion,一种无需微调的模型,可对图像中多个实例的外观和位置进行个性化定制。为提高实例定位精度,引入布局网络(Layout-net),利用显式布局信息与实例区域交叉注意力模块控制生成位置;为提升外观保真度,采用外观网络(appearance-net)提取参考图像中的实例特征,并通过交叉注意力机制融入扩散模型。在COCO和OpenImages数据集上的大量实验表明,该方法在布局与外观引导生成任务中达到当前最优性能。
原文摘要 · Abstract (English)
Recently, text-to-image models based on diffusion have achieved remarkable success in generating high-quality images. However, the challenge of personalized, controllable generation of instances within these images remains an area in need of further development. In this paper, we present LocRef-Diffusion, a novel, tuning-free model capable of personalized customization of multiple instances' appearance and position within an image. To enhance the precision of instance placement, we introduce a Layout-net, which controls instance generation locations by leveraging both explicit instance layout information and an instance region cross-attention module. To improve the appearance fidelity to reference images, we employ an appearance-net that extracts instance appearance features and integrates them into the diffusion model through cross-attention mechanisms. We conducted extensive experiments on the COCO and OpenImages datasets, and the results demonstrate that our proposed method achieves state-of-the-art performance in layout and appearance guided generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。