arXiv:2603.07604cs.CV2026-03

用可学习嵌入替代三平面,实现更高效高质的实时说话头生成

EmbedTalk: Triplane-Free Talking Head Synthesis using Embedding-Driven Gaussian Deformation

  • 用可学习嵌入驱动高斯点变形,取代传统三平面编码
  • 在移动显卡上实测超过60帧/秒,渲染质量与对齐度领先
  • 模型体积更小,适合移动端实时应用

实时说话头合成日益依赖可变形3D高斯溅射(3DGS),因其延迟低。三平面是编码高斯点以进行形变的标准方式,因其提供连续域和明确的空间关系。然而,三平面受网格分辨率限制,并在将三维体数据投影到二维子空间时引入近似误差。近期工作表明,学习到的嵌入在4D场景重建中驱动时序形变具有优势。我们提出EmbedTalk,展示如何利用此类嵌入建模说话头中的语音驱动形变。通过全面实验,我们证明EmbedTalk在渲染质量、唇部同步和运动一致性上优于现有3DGS方法,且与顶尖生成模型性能相当。此外,用学习嵌入替代三平面编码,使模型显著更紧凑,在移动GPU(RTX 2060 6 GB)上实现超过60 FPS。代码将在接受后公开。

原文摘要 · Abstract (English)

Real-time talking head synthesis increasingly relies on deformable 3D Gaussian Splatting (3DGS) due to its low latency. Tri-planes are the standard choice for encoding Gaussians prior to deformation, since they provide a continuous domain with explicit spatial relationships. However, tri-plane representations are limited by grid resolution and approximation errors introduced by projecting 3D volumetric fields onto 2D subspaces. Recent work has shown the superiority of learnt embeddings for driving temporal deformations in 4D scene reconstruction. We introduce $\textbf{EmbedTalk}$, which shows how such embeddings can be leveraged for modelling speech deformations in talking head synthesis. Through comprehensive experiments, we show that EmbedTalk outperforms existing 3DGS-based methods in rendering quality, lip synchronisation, and motion consistency, while remaining competitive with state-of-the-art generative models. Moreover, replacing the tri-plane encoding with learnt embeddings enables significantly more compact models that achieve over 60 FPS on a mobile GPU (RTX 2060 6 GB). Our code will be placed in the public domain on acceptance.

3D高斯说话头生成嵌入驱动实时合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。