让输入人物保持原样,只换背景,生成符合描述的新图。
SceneBooth: Diffusion-based Framework for Subject-preserved Text-to-Image Generation
- 固定人物图像,仅生成匹配文本的背景
- 通过布局与绘画模块,实现高保真人物还原
- 适合需要人物一致性生成的创意设计场景
为满足个性化图像生成需求,基于文本提示生成人物驱动图像的方法受到广泛关注。现有方法通常学习人物表征并将其融入提示嵌入以引导生成,但难以保证人物外观的一致性。本文提出SceneBooth框架,接收人物图像、物体短语和文本提示作为输入,不学习或重建人物,而是固定原始人物图像,仅生成与其协调的背景图像。为此,该框架引入两个关键组件:多模态布局生成模块,根据文本描述、物体短语和人物视觉信息确定人物的位置与尺度;背景绘制模块,将ControlNet与门控自注意力机制集成到潜在扩散模型中,生成与人物和谐匹配的背景。实验表明,SceneBooth在人物保真度、图像融合度和整体质量上显著优于基线方法。
原文摘要 · Abstract (English)
Due to the demand for personalizing image generation, subject-driven text-to-image generation method, which creates novel renditions of an input subject based on text prompts, has received growing research interest. Existing methods often learn subject representation and incorporate it into the prompt embedding to guide image generation, but they struggle with preserving subject fidelity. To solve this issue, this paper approaches a novel framework named SceneBooth for subject-preserved text-to-image generation, which consumes inputs of a subject image, object phrases and text prompts. Instead of learning the subject representation and generating a subject, our SceneBooth fixes the given subject image and generates its background image guided by the text prompts. To this end, our SceneBooth introduces two key components, i.e., a multimodal layout generation module and a background painting module. The former determines the position and scale of the subject by generating appropriate scene layouts that align with text captions, object phrases, and subject visual information. The latter integrates two adapters (ControlNet and Gated Self-Attention) into the latent diffusion model to generate a background that harmonizes with the subject guided by scene layouts and text descriptions. In this manner, our SceneBooth ensures accurate preservation of the subject's appearance in the output. Quantitative and qualitative experimental results demonstrate that SceneBooth significantly outperforms baseline methods in terms of subject preservation, image harmonization and overall quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。