通过融合不同LoRA模型的去噪潜空间,实现传统艺术风格的区域化可控混合。
Style Composition within Distinct LoRA modules for Traditional Art
- 在去噪过程中对多风格专用模型的低噪声潜变量进行区域融合。
- 支持用户指定区域混合多种艺术风格,且不破坏单个风格的完整性。
- 结合深度图控制增强结构一致性,适合艺术创作与跨风格设计。
基于扩散模型的文本到图像生成已能在文本提示下合成多样图像,并通过风格个性化捕捉特定艺术风格。然而,其潜空间耦合及缺乏平滑插值的问题,导致难以在受控、区域化的方式下应用不同绘画技法,常出现一种风格主导的情况。为此,我们提出一种零样本扩散流程,通过在独立训练的风格专精模型的流匹配去噪过程中对去噪潜变量进行风格组合,自然融合多种风格。利用低噪声潜变量携带更强风格信息的特点,采用空间掩码在异构扩散流程间融合潜变量,实现精确的区域风格控制。该机制在保持各风格自身保真度的同时,支持用户引导的混合。此外,为确保不同模型间的结构一致性,我们在扩散框架中引入深度图条件控制(ControlNet)。定性和定量实验表明,该方法能有效根据给定掩码实现区域化的风格混合。
原文摘要 · Abstract (English)
Diffusion-based text-to-image models have achieved remarkable results in synthesizing diverse images from text prompts and can capture specific artistic styles via style personalization. However, their entangled latent space and lack of smooth interpolation make it difficult to apply distinct painting techniques in a controlled, regional manner, often causing one style to dominate. To overcome this, we propose a zero-shot diffusion pipeline that naturally blends multiple styles by performing style composition on the denoised latents predicted during the flow-matching denoising process of separately trained, style-specialized models. We leverage the fact that lower-noise latents carry stronger stylistic information and fuse them across heterogeneous diffusion pipelines using spatial masks, enabling precise, region-specific style control. This mechanism preserves the fidelity of each individual style while allowing user-guided mixing. Furthermore, to ensure structural coherence across different models, we incorporate depth-map conditioning via ControlNet into the diffusion framework. Qualitative and quantitative experiments demonstrate that our method successfully achieves region-specific style mixing according to the given masks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。