arXiv:2409.08520cs.CV2024-09被引 15

让文字生成图像时精准定位主体和背景位置,支持多主体个性化

GroundingBooth: Grounding Text-to-Image Customization

  • 引入新模块与注意力机制,实现零样本实例级空间定位
  • 在布局引导生成与图文定制任务中表现优异,保持身份一致性
  • 适合需要精确控制图像结构的个性化生成场景

近期文本到图像定制方法主要关注输入主体的身份保持,但常无法控制物体的空间位置与大小。我们提出GroundingBooth,可在文本到图像定制任务中实现零样本、实例级的前景主体与背景物体空间定位。所提出的接地模块与主体引导交叉注意力层,使生成图像具备准确的布局对齐、身份保留与强文本-图像一致性。此外,模型可无缝支持多主体个性化。实验表明,该模型在布局引导图像合成与文本到图像定制任务中均表现强劲。项目主页见 https://groundingbooth.github.io。

原文摘要 · Abstract (English)

Recent approaches in text-to-image customization have primarily focused on preserving the identity of the input subject, but often fail to control the spatial location and size of objects. We introduce GroundingBooth, which achieves zero-shot, instance-level spatial grounding on both foreground subjects and background objects in the text-to-image customization task. Our proposed grounding module and subject-grounded cross-attention layer enable the creation of personalized images with accurate layout alignment, identity preservation, and strong text-image coherence. In addition, our model seamlessly supports personalization with multiple subjects. Our model shows strong results in both layout-guided image synthesis and text-to-image customization tasks. The project page is available at https://groundingbooth.github.io.

图像生成空间定位个性化定制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。