用单视频快速生成高保真可驱动的3D头像,兼顾速度与泛化能力
ELITE: Efficient Gaussian Head Avatar from a Monocular Video via Learned Initialization and TEst-time Generative Adaptation
- 结合3D和2D先验优势,通过前馈模型快速初始化高斯头像
- 测试时采用单步扩散增强,恢复细节且合成速度提升60倍
- 适合需要实时生成高质量虚拟形象的场景应用
我们提出ELITE,一种基于单目视频的高效高斯头像生成方法,通过学习初始化与测试时生成适应实现。现有方法依赖3D或2D生成先验来弥补单目视频缺失的视觉线索,但3D先验泛化能力弱,2D先验计算量大且易产生身份幻觉。我们发现二者存在互补性,设计了一个前馈式Mesh2Gaussian先验模型(MGPM),实现高斯头像的快速初始化。为缩小测试时域差距,引入测试时生成适应阶段,利用真实与合成图像联合监督。不同于以往全扩散去噪策略,我们提出基于渲染引导的单步扩散增强器,基于高斯头像渲染结果恢复缺失细节。实验表明,ELITE在复杂表情下仍能生成视觉质量更优的头像,合成速度比2D生成先验方法快60倍。
原文摘要 · Abstract (English)
We introduce ELITE, an Efficient Gaussian head avatar synthesis from a monocular video via Learned Initialization and TEst-time generative adaptation. Prior works rely either on a 3D data prior or a 2D generative prior to compensate for missing visual cues in monocular videos. However, 3D data prior methods often struggle to generalize in-the-wild, while 2D generative prior methods are computationally heavy and prone to identity hallucination. We identify a complementary synergy between these two priors and design an efficient system that achieves high-fidelity animatable avatar synthesis with strong in-the-wild generalization. Specifically, we introduce a feed-forward Mesh2Gaussian Prior Model (MGPM) that enables fast initialization of a Gaussian avatar. To further bridge the domain gap at test time, we design a test-time generative adaptation stage, leveraging both real and synthetic images as supervision. Unlike previous full diffusion denoising strategies that are slow and hallucination-prone, we propose a rendering-guided single-step diffusion enhancer that restores missing visual details, grounded on Gaussian avatar renderings. Our experiments demonstrate that ELITE produces visually superior avatars to prior works, even for challenging expressions, while achieving 60x faster synthesis than the 2D generative prior method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。