arXiv:2512.22602cs.CV2025-12被引 2

让3D虚拟人说话更像本人,精准还原独特语调和口型。

PTalker: Personalized Speech-Driven 3D Talking Head Animation via Style Disentanglement and Modality Alignment

  • 分离语音与表情的风格内容,保留个性化说话习惯。
  • 三重对齐机制提升口型与语音同步精度。
  • 适合需要高个性化的虚拟主播、数字人应用。

语音驱动的3D说话头生成旨在生成与语音精确同步的逼真面部动画。尽管现有方法在唇同步准确性上取得显著进展,但大多忽视了个体独特的说话风格,限制了个性化与真实感。本文提出新框架PTalker,通过风格解耦保留说话风格,并采用三级模态对齐机制提升唇同步精度。具体地,设计解耦约束将驱动音频与运动序列分别编码至风格与内容空间,增强风格表征;通过图注意力网络实现3D网格顶点的空间对齐,利用交叉注意力捕捉时序依赖关系进行时间对齐,并结合top-k双向对比损失与KL散度约束实现特征对齐。在公开数据集上的大量定性与定量实验表明,PTalker能生成高度逼真且具个人风格的3D说话头,性能超越现有最先进方法。源代码与补充视频见PTalker。

原文摘要 · Abstract (English)

Speech-driven 3D talking head generation aims to produce lifelike facial animations precisely synchronized with speech. While considerable progress has been made in achieving high lip-synchronization accuracy, existing methods largely overlook the intricate nuances of individual speaking styles, which limits personalization and realism. In this work, we present a novel framework for personalized 3D talking head animation, namely "PTalker". This framework preserves speaking style through style disentanglement from audio and facial motion sequences and enhances lip-synchronization accuracy through a three-level alignment mechanism between audio and mesh modalities. Specifically, to effectively disentangle style and content, we design disentanglement constraints that encode driven audio and motion sequences into distinct style and content spaces to enhance speaking style representation. To improve lip-synchronization accuracy, we adopt a modality alignment mechanism incorporating three aspects: spatial alignment using Graph Attention Networks to capture vertex connectivity in the 3D mesh structure, temporal alignment using cross-attention to capture and synchronize temporal dependencies, and feature alignment by top-k bidirectional contrastive losses and KL divergence constraints to ensure consistency between speech and mesh modalities. Extensive qualitative and quantitative experiments on public datasets demonstrate that PTalker effectively generates realistic, stylized 3D talking heads that accurately match identity-specific speaking styles, outperforming state-of-the-art methods. The source code and supplementary videos are available at: PTalker.

3D说话头风格解耦语音同步个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。