arXiv:2602.08615cs.CV2026-02International Conf…被引 1

用两张图生成创意组合,帮设计师找灵感。

Inspiration Seeds: Learning Non-Literal Visual Combinations for Generative Exploration

  • 输入两张图像,自动生成视觉关联的新组合。
  • 无需文字提示,直接在视觉空间重组概念。
  • 适合创意初期的模糊探索阶段,效率高。

尽管生成模型在图像合成中已十分强大,但通常针对精心设计的文本提示进行优化,难以支持创意形成前的开放性视觉探索。设计师常从松散关联的视觉参考中汲取灵感,寻找能激发新想法的隐含联系。我们提出 Inspiration Seeds,一种将图像生成从最终执行转向探索性构思的框架。给定两张输入图像,模型可生成多样且视觉连贯的组合,揭示两者间的潜在关系,无需用户指定文本提示。方法为前馈式,训练数据完全通过视觉手段构建:利用 CLIP 稀疏自编码器提取 CLIP 隐空间中的编辑方向并分离概念对。摆脱语言依赖,实现快速直观的重新组合,支持创作早期模糊阶段的视觉构思。

原文摘要 · Abstract (English)

While generative models have become powerful tools for image synthesis, they are typically optimized for executing carefully crafted textual prompts, offering limited support for the open-ended visual exploration that often precedes idea formation. In contrast, designers frequently draw inspiration from loosely connected visual references, seeking emergent connections that spark new ideas. We propose Inspiration Seeds, a generative framework that shifts image generation from final execution to exploratory ideation. Given two input images, our model produces diverse, visually coherent compositions that reveal latent relationships between inputs, without relying on user-specified text prompts. Our approach is feed-forward, trained on synthetic triplets of decomposed visual aspects derived entirely through visual means: we use CLIP Sparse Autoencoders to extract editing directions in CLIP latent space and isolate concept pairs. By removing the reliance on language and enabling fast, intuitive recombination, our method supports visual ideation at the early and ambiguous stages of creative work.

创意生成视觉探索无文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。