用扩散模型生成稀有创意图像,让视觉内容更独特
Creative Image Generation with Diffusion Models
- 基于CLIP空间中图像概率反推创意度,引导生成低概率稀有图像
- 引入回拉机制,在保持图像质量前提下提升创意性
- 适合追求视觉创新的设计师与内容创作者
创意图像生成已成为研究热点,旨在生成突破想象边界的新颖高质量图像。本文提出一种基于扩散模型的创意生成框架,将创意定义为图像在CLIP嵌入空间中存在概率的倒数。与依赖人工概念混合或排除子类别的传统方法不同,本方法计算生成图像的概率分布,并引导其向低概率区域迁移,从而产出罕见、富有想象力且视觉吸引人的结果。我们还引入回拉机制,在不牺牲视觉保真度的前提下实现高创意输出。在文本到图像扩散模型上的大量实验表明,该框架有效且高效,能生成独特、新颖且引人思考的图像。本工作为生成模型中的创造力提供了新视角,为视觉内容合成中的创新提供了系统性方法。
原文摘要 · Abstract (English)
Creative image generation has emerged as a compelling area of research, driven by the need to produce novel and high-quality images that expand the boundaries of imagination. In this work, we propose a novel framework for creative generation using diffusion models, where creativity is associated with the inverse probability of an image's existence in the CLIP embedding space. Unlike prior approaches that rely on a manual blending of concepts or exclusion of subcategories, our method calculates the probability distribution of generated images and drives it towards low-probability regions to produce rare, imaginative, and visually captivating outputs. We also introduce pullback mechanisms, achieving high creativity without sacrificing visual fidelity. Extensive experiments on text-to-image diffusion models demonstrate the effectiveness and efficiency of our creative generation framework, showcasing its ability to produce unique, novel, and thought-provoking images. This work provides a new perspective on creativity in generative models, offering a principled method to foster innovation in visual content synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。