用扩散模型实现360度一致人脸视角生成
SpinMeRound: Consistent Multi-View Identity Generation Using Diffusion Models
- 输入多视角图像与身份嵌入,生成一致人脸
- 实现360度头像合成,优于现有最先进方法
- 适合需要高质量多视角人脸生成的场景
尽管扩散模型取得进展,从新视角生成真实头像仍是重大挑战。现有方法多局限于有限角度,主要关注正面或近正面视角。尽管大型扩散模型在处理3D场景方面表现稳健,但在面部数据上表现不佳,因其结构复杂且易陷入恐怖谷效应。本文提出SpinMeRound,一种基于扩散模型的方法,可生成一致且准确的多视角头像。通过结合多个输入视角与身份嵌入,该方法能有效合成目标人物的多样视角,同时稳健保持其独特身份特征。实验表明,模型具备360度头像合成能力,性能超越当前最先进的多视角扩散模型。
原文摘要 · Abstract (English)
Despite recent progress in diffusion models, generating realistic head portraits from novel viewpoints remains a significant challenge. Most current approaches are constrained to limited angular ranges, predominantly focusing on frontal or near-frontal views. Moreover, although the recent emerging large-scale diffusion models have been proven robust in handling 3D scenes, they underperform on facial data, given their complex structure and the uncanny valley pitfalls. In this paper, we propose SpinMeRound, a diffusion-based approach designed to generate consistent and accurate head portraits from novel viewpoints. By leveraging a number of input views alongside an identity embedding, our method effectively synthesizes diverse viewpoints of a subject whilst robustly maintaining its unique identity features. Through experimentation, we showcase our model's generation capabilities in 360 head synthesis, while beating current state-of-the-art multiview diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。