让AI图像生成同时精准匹配文字描述和特定风格。
StyleBlend: Enhancing Style-Specific Content Creation in Text-to-Image Diffusion Models
- 拆分风格为构图与纹理,分别学习并融合。
- 在不干扰内容生成的前提下实现风格混合。
- 解决文本错位和风格弱的问题,适合艺术创作场景。
在文本到图像扩散模型中,生成既符合文本提示又具有特定艺术风格的视觉上引人注目的图像仍是一大挑战。本文提出StyleBlend,一种从少量参考图像中学习并应用风格表示的方法,实现内容与风格双重对齐。该方法创新性地将风格分解为构图与纹理两个组件,分别采用不同策略进行学习。随后通过两条合成分支,分别聚焦于对应风格成分,利用共享特征实现高效风格融合,且不影响内容生成。StyleBlend有效解决了以往方法中存在的文本对齐偏差与风格表达不足的问题。大量定性和定量对比实验验证了该方法的优越性。
原文摘要 · Abstract (English)
Synthesizing visually impressive images that seamlessly align both text prompts and specific artistic styles remains a significant challenge in Text-to-Image (T2I) diffusion models. This paper introduces StyleBlend, a method designed to learn and apply style representations from a limited set of reference images, enabling content synthesis of both text-aligned and stylistically coherent. Our approach uniquely decomposes style into two components, composition and texture, each learned through different strategies. We then leverage two synthesis branches, each focusing on a corresponding style component, to facilitate effective style blending through shared features without affecting content generation. StyleBlend addresses the common issues of text misalignment and weak style representation that previous methods have struggled with. Extensive qualitative and quantitative comparisons demonstrate the superiority of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。