用多张图片+文字,零训练生成复杂组合图像
IP-Composer: Semantic Composition of Visual Concepts
- 通过文本定位图片中的概念,融合多图特征生成新图
- 无需训练即可实现跨图像概念组合,控制更精准
- 适合需要精细视觉控制的创意设计场景
内容创作者常从多个视觉来源获取灵感,融合不同元素生成新作品。现有计算方法虽能基于文本生成组合图像,但文本难以精确控制视觉细节。基于图像的方法虽能捕捉更细微特征,但通常受限于可处理概念范围,且需昂贵训练或专用数据。我们提出IP-Composer,一种无需训练的组合图像生成方法,可同时利用多张参考图像,并通过自然语言描述每张图中要提取的概念。该方法基于IP-Adapter,将输入图像的CLIP嵌入投影到由文本定义的概念特定子空间中,构建复合嵌入。通过全面评估,结果表明该方法能实现更大范围、更精确的视觉概念组合。
原文摘要 · Abstract (English)
Content creators often draw inspiration from multiple visual sources, combining distinct elements to craft new compositions. Modern computational approaches now aim to emulate this fundamental creative process. Although recent diffusion models excel at text-guided compositional synthesis, text as a medium often lacks precise control over visual details. Image-based composition approaches can capture more nuanced features, but existing methods are typically limited in the range of concepts they can capture, and require expensive training procedures or specialized data. We present IP-Composer, a novel training-free approach for compositional image generation that leverages multiple image references simultaneously, while using natural language to describe the concept to be extracted from each image. Our method builds on IP-Adapter, which synthesizes novel images conditioned on an input image's CLIP embedding. We extend this approach to multiple visual inputs by crafting composite embeddings, stitched from the projections of multiple input images onto concept-specific CLIP-subspaces identified through text. Through comprehensive evaluation, we show that our approach enables more precise control over a larger range of visual concept compositions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。