无需文本提示,融合多图特征生成新图像
Zero-Shot Visual Concept Blending Without Text Guidance
- 利用多参考图在CLIP空间中区分共性与独特特征
- 可灵活转移纹理、形状、风格等抽象属性
- 适合艺术设计等领域实现精准视觉融合
我们提出一种新的零样本图像生成技术——视觉概念融合,可在不依赖额外训练或文本提示的情况下,对多个参考图像的特征进行细粒度控制,将其迁移至源图像。当仅有一个参考图时难以确定具体迁移内容;而使用多个参考图时,该方法通过对比分析,选择性地将共性与独特特征融入生成结果。在部分解耦的CLIP嵌入空间(基于IP-Adapter)中操作,支持纹理、形状、运动、风格及更抽象的概念转换。实验涵盖风格迁移、形态变换和概念融合等多种任务,成功实现笔触风格、流线型、动态感等细微或抽象属性的自然结合。用户研究表明,参与者能准确识别出预期转移的特征。该方法具有高灵活性与高层次控制能力,适用于艺术、设计和内容创作领域。
原文摘要 · Abstract (English)
We propose a novel, zero-shot image generation technique called "Visual Concept Blending" that provides fine-grained control over which features from multiple reference images are transferred to a source image. If only a single reference image is available, it is difficult to isolate which specific elements should be transferred. However, using multiple reference images, the proposed approach distinguishes between common and unique features by selectively incorporating them into a generated output. By operating within a partially disentangled Contrastive Language-Image Pre-training (CLIP) embedding space (from IP-Adapter), our method enables the flexible transfer of texture, shape, motion, style, and more abstract conceptual transformations without requiring additional training or text prompts. We demonstrate its effectiveness across a diverse range of tasks, including style transfer, form metamorphosis, and conceptual transformations, showing how subtle or abstract attributes (e.g., brushstroke style, aerodynamic lines, and dynamism) can be seamlessly combined into a new image. In a user study, participants accurately recognized which features were intended to be transferred. Its simplicity, flexibility, and high-level control make Visual Concept Blending valuable for creative fields such as art, design, and content creation, where combining specific visual qualities from multiple inspirations is crucial.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。