arXiv:2503.06956cs.CV2025-03CVPR被引 15

通过隐式文本融合实现多概念图像生成的高效扩展

LatexBlend: Scaling Multi-concept Customized Generation with Latent Textual Blending

  • 在编码后隐空间融合多个概念,实现灵活组合
  • 支持多概念生成,显著降低去噪偏差,保持布局一致性
  • 适合需要快速生成高保真定制图像的用户

定制化文生图生成基于文本提示将用户指定的概念融入新场景。扩大定制概念数量以满足用户创作需求,但现有方法在生成质量和计算效率上面临挑战。本文提出LaTexBlend框架,有效且高效地扩展多概念定制生成。其核心思想是在文本编码器之后的隐式文本空间中表示单个概念,并融合多个概念,该空间经线性投影生成。LaTexBlend对每个概念独立定制,将其存储于概念库中,使用紧凑的隐式文本特征表示,充分捕捉概念信息以保证高保真度。推理时,可从库中自由无缝组合概念,具备两大优势:1)良好可扩展性;2)显著降低去噪偏差,保持连贯布局。大量实验表明,LaTexBlend能灵活整合多个定制概念,结构和谐、主体保真度高,在生成质量与计算效率上均显著优于基线。代码将公开。

原文摘要 · Abstract (English)

Customized text-to-image generation renders user-specified concepts into novel contexts based on textual prompts. Scaling the number of concepts in customized generation meets a broader demand for user creation, whereas existing methods face challenges with generation quality and computational efficiency. In this paper, we propose LaTexBlend, a novel framework for effectively and efficiently scaling multi-concept customized generation. The core idea of LaTexBlend is to represent single concepts and blend multiple concepts within a Latent Textual space, which is positioned after the text encoder and a linear projection. LaTexBlend customizes each concept individually, storing them in a concept bank with a compact representation of latent textual features that captures sufficient concept information to ensure high fidelity. At inference, concepts from the bank can be freely and seamlessly combined in the latent textual space, offering two key merits for multi-concept generation: 1) excellent scalability, and 2) significant reduction of denoising deviation, preserving coherent layouts. Extensive experiments demonstrate that LaTexBlend can flexibly integrate multiple customized concepts with harmonious structures and high subject fidelity, substantially outperforming baselines in both generation quality and computational efficiency. Our code will be publicly available.

文生图多概念生成隐空间融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。