仅用一张图生成可驱动的逼真3D人脸,支持新身份和表情变化。
SEGA: Drivable 3D Gaussian Head Avatar from a Single Image
- 融合2D与3D先验,构建分层UV空间高斯点云框架
- 单张图像生成高质量3D头像,实时渲染且表情自然
- 适合虚拟社交、数字人创作等需要快速建模的场景
从有限输入生成逼真3D头像在虚拟现实、远程通信和数字娱乐中日益重要。尽管神经渲染和3D高斯溅射技术已实现高质量数字人创建与动画,但多数方法依赖多视角或多图像输入,限制了实际应用。本文提出SEGA——一种基于单张图像的可驱动3D高斯头像生成方法,结合大规模2D数据的通用先验与多视角、多表情、多身份数据学习的3D先验,实现对未见身份的鲁棒泛化,并保证新视角与表情下的3D一致性。我们设计了分层UV空间高斯溅射框架,利用基于FLAME的结构先验,采用双分支架构解耦动态与静态面部成分:动态分支捕捉表情驱动的细节,静态分支聚焦于表情不变区域,实现高效参数推断与预计算。该设计最大化有限3D数据的利用率,实现动画与渲染的实时性能。此外,通过人物特异性微调进一步提升生成头像的保真度与真实感。实验表明,本方法在泛化能力、身份保留与表情真实性方面优于现有最先进方法,推动了一次性头像创建在实际应用中的发展。
原文摘要 · Abstract (English)
Creating photorealistic 3D head avatars from limited input has become increasingly important for applications in virtual reality, telepresence, and digital entertainment. While recent advances like neural rendering and 3D Gaussian splatting have enabled high-quality digital human avatar creation and animation, most methods rely on multiple images or multi-view inputs, limiting their practicality for real-world use. In this paper, we propose SEGA, a novel approach for Single-imagE-based 3D drivable Gaussian head Avatar creation that combines generalized prior models with a new hierarchical UV-space Gaussian Splatting framework. SEGA seamlessly combines priors derived from large-scale 2D datasets with 3D priors learned from multi-view, multi-expression, and multi-ID data, achieving robust generalization to unseen identities while ensuring 3D consistency across novel viewpoints and expressions. We further present a hierarchical UV-space Gaussian Splatting framework that leverages FLAME-based structural priors and employs a dual-branch architecture to disentangle dynamic and static facial components effectively. The dynamic branch encodes expression-driven fine details, while the static branch focuses on expression-invariant regions, enabling efficient parameter inference and precomputation. This design maximizes the utility of limited 3D data and achieves real-time performance for animation and rendering. Additionally, SEGA performs person-specific fine-tuning to further enhance the fidelity and realism of the generated avatars. Experiments show our method outperforms state-of-the-art approaches in generalization ability, identity preservation, and expression realism, advancing one-shot avatar creation for practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。