arXiv:2501.12224cs.CV2025-01被引 59

用单图生成多概念组合,实现精细可控的图像个性化。

TokenVerse: Versatile Multi-concept Personalization in Token Modulation Space

  • 通过调制空间优化,从单图中分离出多个视觉概念。
  • 支持物体、材质、姿态等多类概念组合生成,效果优于现有方法。
  • 适合需要精细控制个性化图像生成的创作者与研究者。

我们提出 TokenVerse——一种基于预训练文生图扩散模型的多概念个性化方法。该框架仅需单张图像即可解耦复杂的视觉元素与属性,并支持从多张图像中提取的概念进行无缝组合生成。与现有方法不同,TokenVerse可处理每张图像包含多个概念的情况,支持对象、配饰、材质、姿态和光照等多种概念。本工作基于DiT架构的文生图模型,发现调制空间具有语义性,能对复杂概念实现局部控制。在此基础上,设计了一种基于优化的框架,输入图像与文本描述后,为每个词在调制空间中寻找独立方向,进而生成组合新图像。实验表明,TokenVerse在挑战性个性化场景下表现优异,显著优于现有方法。

原文摘要 · Abstract (English)

We present TokenVerse -- a method for multi-concept personalization, leveraging a pre-trained text-to-image diffusion model. Our framework can disentangle complex visual elements and attributes from as little as a single image, while enabling seamless plug-and-play generation of combinations of concepts extracted from multiple images. As opposed to existing works, TokenVerse can handle multiple images with multiple concepts each, and supports a wide-range of concepts, including objects, accessories, materials, pose, and lighting. Our work exploits a DiT-based text-to-image model, in which the input text affects the generation through both attention and modulation (shift and scale). We observe that the modulation space is semantic and enables localized control over complex concepts. Building on this insight, we devise an optimization-based framework that takes as input an image and a text description, and finds for each word a distinct direction in the modulation space. These directions can then be used to generate new images that combine the learned concepts in a desired configuration. We demonstrate the effectiveness of TokenVerse in challenging personalization settings, and showcase its advantages over existing methods. project's webpage in https://token-verse.github.io/

图像生成个性化扩散模型多概念

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。