一键生成可动画的高精度3D人脸,支持大视角和动态4D重建。
FA-LAM: Focus-Aware Large Avatar Model for One-Shot 4D Animatable Gaussian Head

- 通过注意力正则化与双阶段训练解耦重建与动画任务。
- 在大视角下仍保持面部细节清晰,支持多视角与流式4D重建。
- 适合需要高质量人脸动画的应用,如虚拟偶像、数字人。
我们提出FA-LAM,一种面向单次输入的可动画高斯人脸大模型,同时实现静态3D与动态4D全头重建。针对现有先进方法中注意力机制误激活及重建与动画任务冲突的问题,我们分析发现两大关键缺陷:(1) 错误且噪声干扰的注意力激活;(2) 重建与动画目标间的内在冲突。为解决第一点,引入对称性与语义感知注意力正则化,利用人脸固有的语义与结构对称性;为解耦任务目标,设计新型双阶段训练流程,将大视角幻觉与动画能力分离至不同模块。此外,通过核心自回归修改与可见性感知标记融合策略,高效实现多视角与流式4D重建。整体上,FA-LAM可在大视角与精细面部区域实现更优的可动画高斯全头重建质量。
原文摘要 · Abstract (English)
We propose FA-LAM, a Focus-Aware Large Avatar Model for one-shot animatable Gaussian head creation, while simultaneously enabling static 3D and dynamic 4D full-head recovery. The core of our method lies in a thorough analysis of the attention mechanisms and the entangled reconstruction and animation training pipeline adopted by prior state-of-the-art approaches. Our analysis identifies two main factors that compromise the quality of 3D full-head generation: (1) incorrect and noisy attention activations, and (2) conflicts between the tasks of reconstruction and animation. To address the first issue, we introduce a symmetric and semantic attention regularization strategy that leverages the inherent semantics and structural symmetry of human heads. To disentangle the objectives of reconstruction and animation, we develop a novel dual-phase training pipeline that separates the model's capabilities for large-view hallucination and animation into distinct modules. Moreover, we enhance our model to support multi-view and streaming 4D reconstruction in an efficient and memory-friendly manner through a core autoregressive modification with tailored visibility-aware token fusion. Collectively, these innovations enable FA-LAM to reconstruct animatable Gaussian full heads with superior quality, particularly in fine facial regions and large viewing angles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。