arXiv:2505.19063cs.CV2025-05被引 2

无需微调,快速生成风格化图像。

Training-free Stylized Text-to-Image Generation with Fast Inference

  • 利用潜在一致性模型的自一致性提取风格统计特征。
  • 通过注意力归一化混合实现风格匹配,生成结果更贴近参考风格。
  • 适合需要快速风格迁移且无训练资源的场景。

尽管扩散模型具有出色的生成能力,但现有的基于这些模型的风格化图像生成方法通常需要使用风格图像进行文本反转或微调,耗时且限制了大规模扩散模型的实际应用。为解决这些问题,我们提出一种新方法 OmniPainter,该方法基于预训练的大规模扩散模型,无需微调或额外优化,即可实现风格化图像生成。具体而言,我们利用潜在一致性模型的自一致性特性,从参考风格图像中提取代表性风格统计信息,用于引导风格化过程。此外,我们引入了自注意力的归一化混合机制,使模型能够从这些统计信息中查询最相关的风格模式,以匹配中间输出内容特征。该机制确保生成结果与参考风格图像的分布高度一致。定性与定量实验表明,所提方法优于当前最优方法。

原文摘要 · Abstract (English)

Although diffusion models exhibit impressive generative capabilities, existing methods for stylized image generation based on these models often require textual inversion or fine-tuning with style images, which is time-consuming and limits the practical applicability of large-scale diffusion models. To address these challenges, we propose a novel stylized image generation method leveraging a pre-trained large-scale diffusion model without requiring fine-tuning or any additional optimization, termed as OmniPainter. Specifically, we exploit the self-consistency property of latent consistency models to extract the representative style statistics from reference style images to guide the stylization process. Additionally, we then introduce the norm mixture of self-attention, which enables the model to query the most relevant style patterns from these statistics for the intermediate output content features. This mechanism also ensures that the stylized results align closely with the distribution of the reference style images. Our qualitative and quantitative experimental results demonstrate that the proposed method outperforms state-of-the-art approaches.

风格迁移扩散模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。