arXiv:2410.04932cs.CV2024-10被引 5

让图像生成支持多模态精细控制,用户可自由指定位置和属性。

OmniBooth: Learning Latent Control for Image Synthesis with Multi-modal Instruction

  • 用高维潜在空间信号统一融合位置、文字和图像条件。
  • 支持实例级开放词汇生成,且可保留个性化身份特征。
  • 适合需要精准布局和自定义风格的图像创作场景。

我们提出OmniBooth,一种支持实例级多模态定制的图像生成框架,实现空间可控生成。用户可通过文本提示或图像参考描述任意物体的位置与属性。给定用户定义的掩码及对应的文本或图像引导,目标是生成图像,使多个对象位于指定坐标,并其属性严格匹配引导信息。该方法显著拓展了文本到图像生成的应用范围,提升了可控性与实用性。核心贡献在于提出的潜在控制信号——一种高维空间特征,能无缝整合空间、文本与图像条件。文本条件扩展ControlNet以支持实例级开放词汇生成;图像条件进一步实现个性化身份的细粒度控制。实际应用中,用户可根据需求灵活选择文本或图像作为多模态引导。大量实验表明,本方法在不同任务与数据集上均展现出更优的图像合成保真度与对齐效果。

原文摘要 · Abstract (English)

We present OmniBooth, an image generation framework that enables spatial control with instance-level multi-modal customization. For all instances, the multimodal instruction can be described through text prompts or image references. Given a set of user-defined masks and associated text or image guidance, our objective is to generate an image, where multiple objects are positioned at specified coordinates and their attributes are precisely aligned with the corresponding guidance. This approach significantly expands the scope of text-to-image generation, and elevates it to a more versatile and practical dimension in controllability. In this paper, our core contribution lies in the proposed latent control signals, a high-dimensional spatial feature that provides a unified representation to integrate the spatial, textual, and image conditions seamlessly. The text condition extends ControlNet to provide instance-level open-vocabulary generation. The image condition further enables fine-grained control with personalized identity. In practice, our method empowers users with more flexibility in controllable generation, as users can choose multi-modal conditions from text or images as needed. Furthermore, thorough experiments demonstrate our enhanced performance in image synthesis fidelity and alignment across different tasks and datasets. Project page: https://len-li.github.io/omnibooth-web/

图像生成多模态可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。