用多视角扩散模型实现真实世界头发迁移,保持视图一致性和高保真度。
Stable-Hair v2: Real-World Hair Transfer via Multiple-View Diffusion Model
- 基于多视角扩散模型,结合极坐标与时间注意力机制实现姿态控制。
- 在三视图数据上训练,生成高质量头像对,迁移细节更真实。
- 适合数字人、虚拟形象等需要多视角一致性的应用开发人员。
尽管基于扩散的方法在捕捉多样且复杂的发型方面表现优异,但其在生成一致且高质量的多视角输出方面仍缺乏研究,而这对数字人和虚拟化身等实际应用至关重要。本文提出 Stable-Hair v2,一种新型基于扩散的多视角头发迁移框架。据我们所知,这是首个利用多视角扩散模型实现鲁棒、高保真、视图一致头发迁移的工作。我们构建了完整的多视角训练数据生成流程,包括基于扩散的光头转换器、数据增强修复模型以及微调后的多视角扩散模型,生成高质量三元组数据:光头图像、参考发型和视角对齐的源-光头配对。我们的多视角头发迁移模型引入极角嵌入进行姿态条件控制,并采用时间注意力层确保视图间过渡平滑。为优化该模型,设计了多阶段训练策略:可控制姿态的潜在身份网络训练、头发提取器训练及时间注意力训练。大量实验表明,该方法能准确将精细真实的发型迁移到目标人物,同时在各视角间实现无缝一致的结果,显著优于现有方法,建立了多视角头发迁移的新基准。代码已公开于 https://github.com/sunkymepro/StableHairV2。
原文摘要 · Abstract (English)
While diffusion-based methods have shown impressive capabilities in capturing diverse and complex hairstyles, their ability to generate consistent and high-quality multi-view outputs -- crucial for real-world applications such as digital humans and virtual avatars -- remains underexplored. In this paper, we propose Stable-Hair v2, a novel diffusion-based multi-view hair transfer framework. To the best of our knowledge, this is the first work to leverage multi-view diffusion models for robust, high-fidelity, and view-consistent hair transfer across multiple perspectives. We introduce a comprehensive multi-view training data generation pipeline comprising a diffusion-based Bald Converter, a data-augment inpainting model, and a face-finetuned multi-view diffusion model to generate high-quality triplet data, including bald images, reference hairstyles, and view-aligned source-bald pairs. Our multi-view hair transfer model integrates polar-azimuth embeddings for pose conditioning and temporal attention layers to ensure smooth transitions between views. To optimize this model, we design a novel multi-stage training strategy consisting of pose-controllable latent IdentityNet training, hair extractor training, and temporal attention training. Extensive experiments demonstrate that our method accurately transfers detailed and realistic hairstyles to source subjects while achieving seamless and consistent results across views, significantly outperforming existing methods and establishing a new benchmark in multi-view hair transfer. Code is publicly available at https://github.com/sunkymepro/StableHairV2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。