仅用少量视频即可快速适配新身份的3D说话头生成方法
Few-Shot Identity Adaptation for 3D Talking Heads via Global Gaussian Field
- 用全局高斯场共享面部结构,实现多身份统一表示
- 仅需少数样本即可完成身份适配,训练效率显著提升
- 适合需要快速部署新身份的交互式应用
基于重建与渲染的3D说话头合成方法虽能保持高质量和强身份保真度,但依赖特定身份模型,每个新身份需从头训练,计算成本高且可扩展性差。为此,我们提出FIAG,一种新型3D说话头合成框架,仅需少量训练视频即可实现高效的身份特异性适配。FIAG引入全局高斯场(Global Gaussian Field),支持在共享场中表示多个身份;同时引入通用运动场(Universal Motion Field),捕捉跨身份的共性运动规律。得益于全局高斯场编码的共享面部结构信息及运动场学习的通用运动先验,本框架可从标准身份表征快速适配至特定身份,仅需极少数据。大量对比与消融实验表明,该方法优于现有最先进方案,验证了其有效性与泛化能力。代码已公开于:https://github.com/gme-hong/FIAG。
原文摘要 · Abstract (English)
Reconstruction and rendering-based talking head synthesis methods achieve high-quality results with strong identity preservation but are limited by their dependence on identity-specific models. Each new identity requires training from scratch, incurring high computational costs and reduced scalability compared to generative model-based approaches. To overcome this limitation, we propose FIAG, a novel 3D speaking head synthesis framework that enables efficient identity-specific adaptation using only a few training footage. FIAG incorporates Global Gaussian Field, which supports the representation of multiple identities within a shared field, and Universal Motion Field, which captures the common motion dynamics across diverse identities. Benefiting from the shared facial structure information encoded in the Global Gaussian Field and the general motion priors learned in the motion field, our framework enables rapid adaptation from canonical identity representations to specific ones with minimal data. Extensive comparative and ablation experiments demonstrate that our method outperforms existing state-of-the-art approaches, validating both the effectiveness and generalizability of the proposed framework. Code is available at: \textit{https://github.com/gme-hong/FIAG}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。