arXiv:2506.24085cs.CVcs.AI2025-06被引 1

用新注意力机制融合图片与文字,生成创意新图像。

Imagine for Me: Creative Conceptual Blending of Real Images and Text via Blended Attention

  • 通过新型混合注意力,分离并融合图文特征。
  • 在真实图像细节保留上优于基线模型。
  • 适合需要创意灵感的设计师或艺术家使用。

将视觉与文本概念融合生成新视觉概念是人类独有的创造力表现,但实际操作中易受认知偏差影响,导致设计陷入局部最优。本文提出T2I扩散模型适配器IT-Blender,可自动化完成跨模态概念融合以增强人类创造力。现有方法在保留真实图像细节或解耦图文输入方面存在不足。IT-Blender利用预训练扩散模型(SD和FLUX)融合干净参考图像与噪声生成图像的潜在表示,并结合新颖的混合注意力机制,在不损失图像细节的前提下,以解耦方式将文本指定物体的视觉概念融入参考图像。实验表明,IT-Blender在跨模态概念融合任务中显著优于基线模型,为图像生成模型在辅助人类创意方面的应用提供了新思路。

原文摘要 · Abstract (English)

Blending visual and textual concepts into a new visual concept is a unique and powerful trait of human beings that can fuel creativity. However, in practice, cross-modal conceptual blending for humans is prone to cognitive biases, like design fixation, which leads to local minima in the design space. In this paper, we propose a T2I diffusion adapter "IT-Blender" that can automate the blending process to enhance human creativity. Prior works related to cross-modal conceptual blending are limited in encoding a real image without loss of details or in disentangling the image and text inputs. To address these gaps, IT-Blender leverages pretrained diffusion models (SD and FLUX) to blend the latent representations of a clean reference image with those of the noisy generated image. Combined with our novel blended attention, IT-Blender encodes the real reference image without loss of details and blends the visual concept with the object specified by the text in a disentangled way. Our experiment results show that IT-Blender outperforms the baselines by a large margin in blending visual and textual concepts, shedding light on the new application of image generative models to augment human creativity.

图像生成概念融合扩散模型创意设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。