arXiv:2411.14793cs.CV2024-11被引 8

通过调整噪声水平提升风格学习能力,让模型更好理解个性化风格。

Style-Friendly SNR Sampler for Style-Driven Generation

  • 在微调时主动提高噪声水平,聚焦风格特征显现阶段。
  • 新方法使模型生成更符合参考图与文本描述的新风格图像。
  • 适合需要定制化风格生成的用户,如艺术创作与个性化内容设计。

近期的文生图扩散模型虽能生成高质量图像,但在学习新个性风格方面表现不佳,限制了独特风格模板的创建。在风格驱动生成中,用户通常提供体现目标风格的参考图像及指定风格属性的文本提示。以往方法多依赖微调,但往往盲目沿用预训练中的目标函数和噪声水平分布,未做适应性调整。我们发现,风格特征主要在较高噪声水平下显现,导致现有微调方法在风格对齐上表现不佳。为此,我们提出风格友好型信噪比采样器(Style-friendly SNR sampler),在微调过程中主动将信噪比(SNR)分布推向更高噪声水平,以聚焦于风格特征显现的关键阶段。该方法显著提升了模型捕捉参考图像与文本提示所指示新风格的能力。实验表明,该方法可有效生成仅靠文本提示难以充分描述的新风格图像,支持个性化内容创作中新型风格模板的构建。

原文摘要 · Abstract (English)

Recent text-to-image diffusion models generate high-quality images but struggle to learn new, personalized styles, which limits the creation of unique style templates. In style-driven generation, users typically supply reference images exemplifying the desired style, together with text prompts that specify desired stylistic attributes. Previous approaches popularly rely on fine-tuning, yet it often blindly utilizes objectives and noise level distributions from pre-training without adaptation. We discover that stylistic features predominantly emerge at higher noise levels, leading current fine-tuning methods to exhibit suboptimal style alignment. We propose the Style-friendly SNR sampler, which aggressively shifts the signal-to-noise ratio (SNR) distribution toward higher noise levels during fine-tuning to focus on noise levels where stylistic features emerge. This enhances models' ability to capture novel styles indicated by reference images and text prompts. We demonstrate improved generation of novel styles that cannot be adequately described solely with a text prompt, enabling the creation of new style templates for personalized content creation.

风格生成扩散模型微调优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。