arXiv:2505.18445cs.CV2025-05NeurIPS被引 30

让扩散模型在换风格时保持细节一致,还能避免风格变差。

OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data

  • 用成对图像训练一致性,让模型学会跨场景保持风格统一。
  • 分两阶段学习,先学风格再保一致,防止风格退化。
  • 插件式设计,兼容任意风格LoRA,适合想稳定出图的用户。

扩散模型显著推动了图像风格化发展,但仍面临两大挑战:(1)复杂场景中保持身份、构图和细节点的一致性;(2)在使用风格LoRA的图像到图像流水线中防止风格退化。GPT-4o展现出极高的风格一致性,凸显开源方法与商用模型之间的差距。为此,我们提出 extbf{OmniConsistency},一种通用一致性插件,基于大规模扩散Transformer(DiTs)。其贡献包括:(1)基于对齐图像对的上下文一致性学习框架,实现鲁棒泛化;(2)两阶段渐进式学习策略,解耦风格学习与一致性保留,缓解风格退化;(3)完全即插即用设计,兼容任意风格LoRA,适用于Flux框架。大量实验表明,OmniConsistency显著提升视觉连贯性和美学质量,性能接近商业顶级模型GPT-4o。

原文摘要 · Abstract (English)

Diffusion models have advanced image stylization significantly, yet two core challenges persist: (1) maintaining consistent stylization in complex scenes, particularly identity, composition, and fine details, and (2) preventing style degradation in image-to-image pipelines with style LoRAs. GPT-4o's exceptional stylization consistency highlights the performance gap between open-source methods and proprietary models. To bridge this gap, we propose \textbf{OmniConsistency}, a universal consistency plugin leveraging large-scale Diffusion Transformers (DiTs). OmniConsistency contributes: (1) an in-context consistency learning framework trained on aligned image pairs for robust generalization; (2) a two-stage progressive learning strategy decoupling style learning from consistency preservation to mitigate style degradation; and (3) a fully plug-and-play design compatible with arbitrary style LoRAs under the Flux framework. Extensive experiments show that OmniConsistency significantly enhances visual coherence and aesthetic quality, achieving performance comparable to commercial state-of-the-art model GPT-4o.

风格迁移扩散模型一致性LoRA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。