统一框架解决风格迁移中内容与风格混淆问题,提升生成质量与稳定性。
UniCSG: Unified High-Fidelity Content-Constrained Style-Driven Generation via Staged Semantic and Frequency Disentanglement
- 分阶段训练:先分离内容与风格的潜在语义,再通过多尺度频率监督修复细节
- 在文本和参考图引导下均实现更高内容保真度与风格一致性
- 适合需要精准风格迁移的图像生成任务,如艺术创作与设计辅助
风格迁移需在匹配目标风格的同时保持内容语义。基于DiT的扩散模型常因内容-风格纠缠导致参考内容泄露和生成不稳定。我们提出UniCSG,一个统一的框架,支持文本引导和参考引导下的内容约束型风格驱动生成。该方法采用分阶段训练:(i) 潜在空间语义解耦阶段,结合低频预处理与条件扰动,促进内容与风格分离;(ii) 潜在空间频率感知细节重建阶段,通过多尺度频率监督精细化生成。此外,引入像素空间奖励学习,使潜在目标与解码后的感知质量对齐。实验表明,该方法在两种设置下均显著提升了内容忠实度、风格匹配度与鲁棒性。
原文摘要 · Abstract (English)
Style transfer must match a target style while preserving content semantics. DiT-based diffusion models often suffer from content-style entanglement, leading to reference-content leakage and unstable generation. We present UniCSG, a unified framework for content-constrained, style-driven generation in both text-guided and reference-guided settings. UniCSG employs staged training: (i) a latent-space semantic disentanglement stage that combines low-frequency preprocessing with conditioning corruption to encourage content-style separation, and (ii) a latent-space frequency-aware detail reconstruction stage that refines details via multi-scale frequency supervision. We further incorporate pixel-space reward learning to align latent objectives with perceptual quality after decoding. Experiments demonstrate improved content faithfulness, style alignment, and robustness in both settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。