arXiv:2512.12963cs.CV2025-12中稿 · WACV 2026

让扩散模型更真实地迁移风格,速度翻倍且不依赖复杂优化

SCAdapter: Content-Style Disentanglement for Diffusion Style Transfer

  • 用CLIP空间分离内容与风格特征,精准提取纯内容和纯风格
  • 多风格融合更快更准,生成图像细节更丰富、风格更忠实
  • 无需反演和推理优化,速度超现有方法2倍,适合实际应用

扩散模型已成为风格迁移的主流方法,但难以实现照片级真实感迁移,常生成类似绘画的效果或丢失细节。现有方法未能有效消除原始内容风格及参考图像内容对迁移结果的干扰。本文提出SCAdapter,利用CLIP图像空间实现内容与风格的有效解耦与融合。核心创新在于系统性地从内容图像中提取纯净内容特征,从风格参考中提取纯粹风格元素,确保迁移的真实性和准确性。该方法通过三个组件增强:可控制风格自适应实例归一化(CSAdaIN)实现多风格精准融合,基于键值存储的注入机制(KVS Injection)实现定向风格整合,以及风格一致性目标保证过程连贯性。大量实验表明,SCAdapter显著优于现有先进方法,涵盖传统与扩散基线。相比其他扩散方法,本方法无需DDIM反演和推理阶段优化,推理速度至少提升2倍,兼具更高效率与更强效果。

原文摘要 · Abstract (English)

Diffusion models have emerged as the leading approach for style transfer, yet they struggle with photo-realistic transfers, often producing painting-like results or missing detailed stylistic elements. Current methods inadequately address unwanted influence from original content styles and style reference content features. We introduce SCAdapter, a novel technique leveraging CLIP image space to effectively separate and integrate content and style features. Our key innovation systematically extracts pure content from content images and style elements from style references, ensuring authentic transfers. This approach is enhanced through three components: Controllable Style Adaptive Instance Normalization (CSAdaIN) for precise multi-style blending, KVS Injection for targeted style integration, and a style transfer consistency objective maintaining process coherence. Comprehensive experiments demonstrate SCAdapter significantly outperforms state-of-the-art methods in both conventional and diffusion-based baselines. By eliminating DDIM inversion and inference-stage optimization, our method achieves at least $2\times$ faster inference than other diffusion-based approaches, making it both more effective and efficient for practical applications.

风格迁移扩散模型解耦高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。