arXiv:2608.05218cs.AIcs.SD2026-08中稿 · ACM MM 2026

用音素驱动3D高斯点云,让虚拟人脸说话更精准

PD-GS: Phoneme-Driven 3DGS for Audio-Driven Talking Heads

论文配图:PD-GS: Phoneme-Driven 3DGS for Audio-Driven Talking Heads
图 1 · 摘自论文原文
  • 引入音素标签与音频上下文融合,提升发音关键帧的准确性
  • 在HDTF数据集上唇部几何误差低至2.66,显著减少漏嘴现象
  • 适合需要高逼真度语音驱动人脸生成的研究与应用

3D高斯点云(3DGS)能快速生成逼真说话头像,但唇部动作常因过度平滑而违背发音约束,如双唇闭合,导致'漏嘴'伪影。主要难点在于:从连续声学嵌入中推断短暂离散的发音事件,回归目标会偏向平均口型。尽管现代自监督语音编码器提供丰富语调和音素信息,却缺乏对齐帧的显式语言目标来清晰区分闭合事件。我们提出音素驱动高斯点云(PD-GS),通过自动语音识别与强制对齐流程获取时间对齐的音素标记,增强3DGS说话人模型。核心组件语言融合模块(LFM)通过可学习门控机制,自适应融合连续音频上下文与离散音素嵌入,既保留音频驱动的流畅性,又强化音素在发音关键段的引导。PD-GS仅使用单目视频训练,依赖图像重建与唇部关键点监督。在HDTF数据集上,其唇部几何表现优于对比基线(LMD 2.66),在复杂音素序列中显著减少闭合违规,生成更具语言准确性的神经化身。

原文摘要 · Abstract (English)

3D Gaussian Splatting (3DGS) enables fast, photorealistic talking-head rendering, yet accurate lip articulation remains elusive: mouth motion is often over-smoothed and may violate hard articulatory constraints such as bilabial closures, producing the notorious ``leaky mouth'' artifact. A key difficulty is that brief, discrete articulatory events are inferred from a continuous acoustic embedding under a regression objective, which biases predictions toward averaged mouth configurations. While modern self-supervised speech encoders provide rich prosodic and phonetic cues, they do not provide an explicit, frame-aligned linguistic target that reliably disambiguates closure-level events. We propose \textbf{Phoneme-Driven Gaussian Splatting (PD-GS)}, which augments a 3DGS talker with time-aligned phoneme tokens obtained from an automatic ASR and forced-alignment pipeline. Our core component, the \textbf{Linguistic Fusion Module (LFM)}, adaptively fuses continuous audio context with discrete phoneme embeddings through a learned gate, allowing the model to preserve smooth audio-driven dynamics while strengthening phoneme guidance on articulation-critical segments. PD-GS is trained purely from monocular video using image reconstruction and lip landmark supervision. On HDTF, PD-GS achieves the best lip geometry among the compared baselines (LMD 2.66) and qualitatively reduces closure violations in challenging phoneme sequences, yielding more linguistically faithful neural avatars.

3D高斯语音驱动唇动生成音素对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。