无需训练即可保持人物身份的风格化图像生成方法
Training-Free Diffusion Framework for Stylized Image Generation with Identity Preservation
- 提出无训练内容一致性损失,聚焦原图细节增强
- 引入马赛克修复内容图像技术,复杂场景下身份保留更佳
- 适合广告营销等需保真人身份的风格化应用
尽管扩散模型展现出强大的生成能力,现有风格迁移技术在实现高质量风格化时往往难以保持人物身份。这一问题在广告营销等实际应用中尤为突出,当人物距离较远或处于群体中时,身份信息易丢失。为此,本文提出一种无需训练的身份保持风格化图像合成框架。核心贡献包括:1)“马赛克修复内容图像”技术,显著提升复杂场景中的身份保留能力;2)无训练内容一致性损失,通过强化对原图的关注,改善细粒度细节的保留。实验表明,该方法在不需模型重训练或微调的前提下,显著优于基线模型,在高风格保真度与强身份完整性之间实现良好平衡。
原文摘要 · Abstract (English)
Although diffusion models have demonstrated remarkable generative capabilities, existing style transfer techniques often struggle to maintain identity while achieving high-quality stylization. This limitation becomes particularly critical in practical applications such as advertising and marketing, where preserving the identity of featured individuals is essential for a campaign's effectiveness. It is particularly severe when subjects are distant from the camera or appear within a group, frequently leading to a significant loss of identity. To address this issue, we introduce a novel, training-free framework for identity-preserved stylized image synthesis. Key contributions include the "Mosaic Restored Content Image" technique, which significantly enhances identity retention in complex scenes, and a training-free content consistency loss that improves the preservation of fine-grained details by directing more attention to the original image during stylization. Our experiments reveal that the proposed approach substantially exceeds the baseline model in concurrently maintaining high stylistic fidelity and robust identity integrity, all without necessitating model retraining or fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。