arXiv:2509.23643cs.CV2025-09被引 1

用图片引导布局,无需训练即可精准控制图像元素位置

Griffin: Generative Reference and Layout Guided Image Composition

  • 以图像为参考,通过单图指定内容与位置,实现无训练控制
  • 支持对象级和部件级组合,可灵活完成多图像布局任务
  • 适合需要精细布局控制的设计师或内容创作者

文本到图像模型已达到高度逼真的生成水平,但在需要更明确指导时,纯文本控制成为瓶颈。精确指定内容及其在图像中的位置对于实现更精细的控制至关重要。本文解决多图像布局控制问题:通过图像而非文本指定所需内容,并指导模型如何放置每个元素。所提方法无需训练,每类参考仅需一张图像,提供显式且简单的对象与部件级组合控制。我们在多种图像合成任务中验证了其有效性。

原文摘要 · Abstract (English)

Text-to-image models have achieved a level of realism that enables the generation of highly convincing images. However, text-based control can be a limiting factor when more explicit guidance is needed. Defining both the content and its precise placement within an image is crucial for achieving finer control. In this work, we address the challenge of multi-image layout control, where the desired content is specified through images rather than text, and the model is guided on where to place each element. Our approach is training-free, requires a single image per reference, and provides explicit and simple control for object and part-level composition. We demonstrate its effectiveness across various image composition tasks.

图像生成布局控制参考引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。