arXiv:2508.14449cs.CV2025-08

用双分支变形场分离通用与个性特征,少样本下实现高保真说话头合成。

D^3-Talker: Dual-Branch Decoupled Deformation Fields for Few-Shot 3D Talking Head Synthesis

  • 分离静态属性场与音频/表情驱动的动态变形场,解耦通用与个性化信息。
  • 仅用少量帧训练即达到领先水平的唇音同步与图像质量。
  • 适合需要快速定制虚拟人形象的研究者与开发者使用。

3D说话头合成的关键挑战在于:为每个目标身份从零训练新模型需依赖长时视频。现有方法虽通过预训练模型提取音频通用特征,但因音频包含无关唇动信息,导致在仅用少数帧训练时难以生成逼真唇部动作,造成唇音不同步和图像质量差。本文提出D^3-Talker,构建静态3D高斯属性场,并分别用音频和面部运动信号控制两个独立的高斯属性变形场,有效解耦通用与个性化变形预测。设计新颖的相似性对比损失函数,在预训练中实现更彻底的解耦。此外,引入粗到精模块以优化渲染图像,缓解头部运动引起的模糊问题,提升整体画质。大量实验表明,该方法在有限训练数据下,于高保真渲染和准确唇音同步方面均优于当前最优方法。代码将在录用后公开。

原文摘要 · Abstract (English)

A key challenge in 3D talking head synthesis lies in the reliance on a long-duration talking head video to train a new model for each target identity from scratch. Recent methods have attempted to address this issue by extracting general features from audio through pre-training models. However, since audio contains information irrelevant to lip motion, existing approaches typically struggle to map the given audio to realistic lip behaviors in the target face when trained on only a few frames, causing poor lip synchronization and talking head image quality. This paper proposes D^3-Talker, a novel approach that constructs a static 3D Gaussian attribute field and employs audio and Facial Motion signals to independently control two distinct Gaussian attribute deformation fields, effectively decoupling the predictions of general and personalized deformations. We design a novel similarity contrastive loss function during pre-training to achieve more thorough decoupling. Furthermore, we integrate a Coarse-to-Fine module to refine the rendered images, alleviating blurriness caused by head movements and enhancing overall image quality. Extensive experiments demonstrate that D^3-Talker outperforms state-of-the-art methods in both high-fidelity rendering and accurate audio-lip synchronization with limited training data. Our code will be provided upon acceptance.

3D说话头少样本学习高斯建模唇音同步

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。