仅用一张人像图,就能重建高保真3D人脸,支持各种角度和表情。
High-Quality 3D Head Reconstruction from Any Single Portrait Image
- 用多视角扩散模型融合身份与表情信息,提升生成一致性。
- 构建21,792帧的高质量人脸数据集,覆盖96个视角和多样表情。
- 可生成96帧环绕视频,适用于影视、虚拟人等高精度需求场景。
本文提出一种从任意单张人像图重建高保真3D头像的方法,不受视角、表情或配饰影响。尽管已有大量工作尝试将2D生成模型用于新视角合成与3D优化,但多数方法在生成高质量3D头像方面仍受限,主要因缺少身份、表情、发型及配饰等关键信息。为此,我们构建了一个新数据集,包含227个数字人肖像序列,共21,792帧,覆盖96种不同视角,涵盖多样化表情与配饰。为提升性能,我们在多视角扩散过程中融入身份与表情信息,通过身份与表情感知的引导与监督,提取精确的面部表征,指导模型生成并强制目标函数以确保生成过程中的身份与表情一致性。最终生成包含96个视图的环绕视频,可用于3D头像模型重建。实验表明,该方法在侧脸视角和复杂配饰等挑战性场景下仍表现稳健。
原文摘要 · Abstract (English)
In this work, we introduce a novel high-fidelity 3D head reconstruction method from a single portrait image, regardless of perspective, expression, or accessories. Despite significant efforts in adapting 2D generative models for novel view synthesis and 3D optimization, most methods struggle to produce high-quality 3D portraits. The lack of crucial information, such as identity, expression, hair, and accessories, limits these approaches in generating realistic 3D head models. To address these challenges, we construct a new high-quality dataset containing 227 sequences of digital human portraits captured from 96 different perspectives, totalling 21,792 frames, featuring diverse expressions and accessories. To further improve performance, we integrate identity and expression information into the multi-view diffusion process to enhance facial consistency across views. Specifically, we apply identity- and expression-aware guidance and supervision to extract accurate facial representations, which guide the model and enforce objective functions to ensure high identity and expression consistency during generation. Finally, we generate an orbital video around the portrait consisting of 96 multi-view frames, which can be used for 3D portrait model reconstruction. Our method demonstrates robust performance across challenging scenarios, including side-face angles and complex accessories
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。