仅用一张图生成高保真可动3D人脸,解决形状不准与身份不一致问题
FaceCraft4D: Animated 3D Facial Avatar Generation from a Single Image
- 通过3D-GAN反演获初始形状,结合图像扩散模型提升多视角纹理一致性
- 引入视频先验和同步驱动信号,实现跨视角表情动画连续自然
- 提出一致-不一致训练策略,有效应对数据不一致带来的重建挑战
我们提出一种新框架,仅需单张图像即可生成高质量、可动画化的4D人脸虚拟形象。尽管近期研究在4D avatar生成方面取得进展,现有方法仍依赖大量多视角数据,或在形状准确性和身份一致性上表现不佳。为此,我们设计了一套综合系统,融合形状、图像和视频先验,生成全视角可动画化的人脸。首先通过3D-GAN反演获得初始粗略形状;随后利用深度引导的图像扭曲信号,结合图像扩散模型增强多视角纹理,确保跨视角一致性;为实现表情动画,引入带有同步驱动信号的视频先验,保证多视角间表达连贯性。此外,我们提出一致-不一致训练机制,有效处理4D重建中的数据不一致性问题。实验表明,该方法在视觉质量上优于现有技术,且在不同视角和表情下保持高度一致性。
原文摘要 · Abstract (English)
We present a novel framework for generating high-quality, animatable 4D avatar from a single image. While recent advances have shown promising results in 4D avatar creation, existing methods either require extensive multiview data or struggle with shape accuracy and identity consistency. To address these limitations, we propose a comprehensive system that leverages shape, image, and video priors to create full-view, animatable avatars. Our approach first obtains initial coarse shape through 3D-GAN inversion. Then, it enhances multiview textures using depth-guided warping signals for cross-view consistency with the help of the image diffusion model. To handle expression animation, we incorporate a video prior with synchronized driving signals across viewpoints. We further introduce a Consistent-Inconsistent training to effectively handle data inconsistencies during 4D reconstruction. Experimental results demonstrate that our method achieves superior quality compared to the prior art, while maintaining consistency across different viewpoints and expressions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。