arXiv:2512.10939cs.CV2025-12被引 3

用音频驱动高斯点云生成稳定实时的3D说话头像

GaussianHeadTalk: Wobble-Free 3D Talking Heads with Audio Driven Gaussian Splatting

  • 结合3D可变形模型与高斯点云,实现个性化头像建模
  • 通过音频直接预测参数,确保时间一致性,消除抖动
  • 单目视频+独立音频输入,支持实时生成,效果媲美主流方法

语音驱动的3D说话头像近期兴起,可用于交互式虚拟形象。但实际应用受限于现有方法:虽视觉保真度高,却存在速度慢或快但时序不稳的问题。扩散模型生成逼真图像,但在单次生成场景下表现不佳。高斯点云方法可实时运行,但面部追踪不准或高斯映射不一致会导致输出不稳定及视频伪影,影响真实场景使用。本文提出基于3D可变形模型(3DMM)的高斯点云映射方法,构建个性化头像;引入基于Transformer的参数预测模块,直接从音频输入驱动,确保时序一致性;仅需单目视频与独立音频,即可实时生成高质量说话头像视频,在定量与定性评估上均达到竞争力水平。

原文摘要 · Abstract (English)

Speech-driven talking heads have recently emerged and enable interactive avatars. However, real-world applications are limited, as current methods achieve high visual fidelity but slow or fast yet temporally unstable. Diffusion methods provide realistic image generation, yet struggle with oneshot settings. Gaussian Splatting approaches are real-time, yet inaccuracies in facial tracking, or inconsistent Gaussian mappings, lead to unstable outputs and video artifacts that are detrimental to realistic use cases. We address this problem by mapping Gaussian Splatting using 3D Morphable Models to generate person-specific avatars. We introduce transformer-based prediction of model parameters, directly from audio, to drive temporal consistency. From monocular video and independent audio speech inputs, our method enables generation of real-time talking head videos where we report competitive quantitative and qualitative performance.

3D说话头高斯点云音频驱动实时生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。