arXiv:2503.23538cs.CV2025-03CVPR被引 17

不增加计算成本,让Stable Diffusion生成更创意的图像。

Enhancing Creative Generation on Stable Diffusion-based Models

  • 通过在去噪过程选择性增强特征提升创意
  • 提供可操作的放大因子选择指南
  • 适用于多种Stable Diffusion模型,无需重新训练

近期文本到图像生成模型,尤其是Stable Diffusion及其压缩变体,已实现高保真度与强文本-图像对齐。然而其创造性仍受限,仅在提示中加入'creative'通常无法获得理想结果。本文提出C3(Creative Concept Catalyst),一种无需训练的方案,旨在增强基于Stable Diffusion模型的创造力。C3通过在去噪过程中选择性放大特征,促进更具创意的输出。我们基于创造力的两个核心方面,提供了放大因子的选择实践指南。C3是首个在扩散模型中无需大量计算成本即可提升创造力的研究。我们在多种Stable Diffusion基线模型上验证了其有效性。

原文摘要 · Abstract (English)

Recent text-to-image generative models, particularly Stable Diffusion and its distilled variants, have achieved impressive fidelity and strong text-image alignment. However, their creative capability remains constrained, as including `creative' in prompts seldom yields the desired results. This paper introduces C3 (Creative Concept Catalyst), a training-free approach designed to enhance creativity in Stable Diffusion-based models. C3 selectively amplifies features during the denoising process to foster more creative outputs. We offer practical guidelines for choosing amplification factors based on two main aspects of creativity. C3 is the first study to enhance creativity in diffusion models without extensive computational costs. We demonstrate its effectiveness across various Stable Diffusion-based models.

图像生成扩散模型创意增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。