仅用一张图生成逼真会动的人脸,且能泛化到陌生人
SDTalk: Structured Facial Priors and Dual-Branch Motion Fields for Generalizable Gaussian Talking Head Synthesis

- 用结构化面部先验+双分支运动场,从单张图重建完整人脸
- 在1080p下实现90帧/秒实时生成,比现有方法更清晰更同步
- 无需微调就能合成没见过的人脸,适合快速生成应用
高质量、实时的说话头生成仍是计算机视觉中的核心挑战。现有基于重建与渲染的方法通常依赖特定身份模型,限制了跨身份泛化能力。为此,我们提出SDTalk,一种基于3D高斯溅射(3DGS)的一次性框架,可在不进行个性化训练或微调的情况下泛化至未见身份。该框架包含两个模块,采用两阶段训练策略:第一阶段将结构化面部先验融入重建模块,并分别预测可见与遮挡区域的3DGS参数,实现仅凭单张图像完成完整头部重建;第二阶段引入双分支运动场,建模粗粒度与细粒度面部动态,提升细节保真度与唇形同步精度。实验表明,SDTalk在视觉质量与推理效率上均优于现有方法。
原文摘要 · Abstract (English)
High-quality, real-time talking head synthesis remains a fundamental challenge in computer vision. Existing reconstruction- and rendering-based methods typically rely on identity-specific models, limiting cross-identity generalization. To address this issue, we propose SDTalk, a one-shot 3D Gaussian Splatting (3DGS)-based framework that generalizes to unseen identities without personalized training or fine-tuning. Our framework comprises two modules with a two-stage training strategy. In the first stage, we incorporate structured facial priors into the reconstruction module and separately predict 3DGS parameters for visible and occluded regions, enabling complete head reconstruction from a single image. In the second stage, we introduce a dual-branch motion field to model coarse and fine facial dynamics, improving detail fidelity and lip synchronization. Experiments demonstrate that SDTalk surpasses existing methods in both visual quality and inference efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。