无需训练即可按布局生成一致角色图像,适合漫画等需要统一人物的场景。
SpotActor: Training-Free Layout-Controlled Consistent Image Generation

- 通过双语义-隐空间优化,实现布局约束下的反向调整
- 生成图像在布局对齐、角色一致性上显著优于现有方法
- 适用于漫画创作等需保持角色一致性的实际应用
文本到图像扩散模型在艺术创作中大幅提升效率,但在漫画制作等典型场景下,难以将每个主体放置到预期位置,也难以保持主体跨图像的一致性。为此,我们提出新任务——布局到一致图像(L2CI)生成,根据给定布局和文本提示生成具有一致性和组合性的图像。为此,我们提出一种新的双重能量引导形式化,在双语义-隐空间中进行优化,构建无需训练的SpotActor框架,包含布局条件反向更新阶段与一致性正向采样阶段。反向阶段设计精细布局能量函数,以类Sigmoid目标模拟注意力激活;正向阶段引入区域互联自注意力(RISA)和语义融合交叉注意力(SFCA),实现跨图像的相互作用。为评估性能,我们构建了基于目标检测数据集的百组合理提示-框对组成的ActorBench基准。大量实验表明,SpotActor有效达成任务要求,在布局对齐、主体一致性、提示符合度及背景多样性方面表现优异,具备实用潜力。
原文摘要 · Abstract (English)
Text-to-image diffusion models significantly enhance the efficiency of artistic creation with high-fidelity image generation. However, in typical application scenarios like comic book production, they can neither place each subject into its expected spot nor maintain the consistent appearance of each subject across images. For these issues, we pioneer a novel task, Layout-to-Consistent-Image (L2CI) generation, which produces consistent and compositional images in accordance with the given layout conditions and text prompts. To accomplish this challenging task, we present a new formalization of dual energy guidance with optimization in a dual semantic-latent space and thus propose a training-free pipeline, SpotActor, which features a layout-conditioned backward update stage and a consistent forward sampling stage. In the backward stage, we innovate a nuanced layout energy function to mimic the attention activations with a sigmoid-like objective. While in the forward stage, we design Regional Interconnection Self-Attention (RISA) and Semantic Fusion Cross-Attention (SFCA) mechanisms that allow mutual interactions across images. To evaluate the performance, we present ActorBench, a specified benchmark with hundreds of reasonable prompt-box pairs stemming from object detection datasets. Comprehensive experiments are conducted to demonstrate the effectiveness of our method. The results prove that SpotActor fulfills the expectations of this task and showcases the potential for practical applications with superior layout alignment, subject consistency, prompt conformity and background diversity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。