用通用先验+个性适配,生成逼真且能大角度转动的说话头像。
GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation
- 分两阶段学习通用3D先验与个体特征,提升泛化能力。
- 在大角度旋转和陌生语音下仍保持高保真与对齐精度。
- 无需重复训练,适合快速生成不同人物的视频。
高质量、可泛化的语音驱动3D说话头像生成仍是难题。现有方法在固定视角和小范围语音变化下表现良好,但在大幅头部旋转和分布外(OOD)语音下性能下降,且需耗时的个性化训练。核心问题在于缺乏足够的3D先验,限制了生成结果的外推能力。为此,我们提出GGTalker,结合通用先验与身份特异性适应实现说话头像合成。采用两阶段先验-适配训练策略,学习高斯头像先验,并适配个体特征。训练音频-表情和表情-视觉先验以捕捉唇部运动普遍规律及头部纹理整体分布。在定制化适配阶段,精确建模个体说话风格与纹理细节。此外,引入颜色MLP生成与运动对齐的精细纹理,使用身体修复器融合渲染结果与背景,生成难以分辨的逼真视频帧。大量实验表明,GGTalker在渲染质量、3D一致性、唇音同步准确率和训练效率方面均达到当前最优水平。
原文摘要 · Abstract (English)
Creating high-quality, generalizable speech-driven 3D talking heads remains a persistent challenge. Previous methods achieve satisfactory results for fixed viewpoints and small-scale audio variations, but they struggle with large head rotations and out-of-distribution (OOD) audio. Moreover, they are constrained by the need for time-consuming, identity-specific training. We believe the core issue lies in the lack of sufficient 3D priors, which limits the extrapolation capabilities of synthesized talking heads. To address this, we propose GGTalker, which synthesizes talking heads through a combination of generalizable priors and identity-specific adaptation. We introduce a two-stage Prior-Adaptation training strategy to learn Gaussian head priors and adapt to individual characteristics. We train Audio-Expression and Expression-Visual priors to capture the universal patterns of lip movements and the general distribution of head textures. During the Customized Adaptation, individual speaking styles and texture details are precisely modeled. Additionally, we introduce a color MLP to generate fine-grained, motion-aligned textures and a Body Inpainter to blend rendered results with the background, producing indistinguishable, photorealistic video frames. Comprehensive experiments show that GGTalker achieves state-of-the-art performance in rendering quality, 3D consistency, lip-sync accuracy, and training efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。