arXiv:2506.23630cs.CV2025-06

无需训练即可让扩散模型融合多个概念生成新图像。

Blending Concepts with Text-to-Image Diffusion Models

  • 通过提示调度、嵌入插值等方法实现概念融合。
  • 100人用户研究显示不同方法在不同条件下表现各异。
  • 适合想探索创意组合的视觉设计与艺术创作人群。

近年来,扩散模型显著提升了文本到图像生成的能力,能轻松将抽象概念转化为高保真图像。本文探讨这些模型是否能在零样本框架下融合不同概念,从具体物体到无形想法,生成连贯的新视觉实体。概念融合旨在将多个概念(以文本提示表达)的关键属性合并为一张新图像,体现各概念的本质。我们测试了四种融合方法,分别利用扩散流程的不同环节(如提示调度、嵌入插值或层间条件控制)。在多样概念类别上的系统实验表明,现代扩散模型确实在不需额外训练或微调的情况下展现出创造性的融合能力。涵盖具体物体融合、复合词生成、艺术风格迁移及建筑地标融合等任务。一项包含100名参与者的广泛用户研究发现,没有单一方法在所有场景中占优:每种技术在特定条件下表现更佳,且提示顺序、概念距离和随机种子等因素会影响结果。这些发现凸显了扩散模型强大的组合潜力,也揭示其对输入微小变化的高度敏感性。

原文摘要 · Abstract (English)

Diffusion models have dramatically advanced text-to-image generation in recent years, translating abstract concepts into high-fidelity images with remarkable ease. In this work, we examine whether they can also blend distinct concepts, ranging from concrete objects to intangible ideas, into coherent new visual entities under a zero-shot framework. Specifically, concept blending merges the key attributes of multiple concepts (expressed as textual prompts) into a single, novel image that captures the essence of each concept. We investigate four blending methods, each exploiting different aspects of the diffusion pipeline (e.g., prompt scheduling, embedding interpolation, or layer-wise conditioning). Through systematic experimentation across diverse concept categories, such as merging concrete concepts, synthesizing compound words, transferring artistic styles, and blending architectural landmarks, we show that modern diffusion models indeed exhibit creative blending capabilities without further training or fine-tuning. Our extensive user study, involving 100 participants, reveals that no single approach dominates in all scenarios: each blending technique excels under certain conditions, with factors like prompt ordering, conceptual distance, and random seed affecting the outcome. These findings highlight the remarkable compositional potential of diffusion models while exposing their sensitivity to seemingly minor input variations.

文本生成图像概念融合扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。