用语音驱动3D高斯头像,实时生成逼真表情动作。
GaussianSpeech: Audio-Driven Gaussian Avatars
- 结合语音与3D高斯溅射,建模面部细微动态。
- 实现真实表情变化,支持多样人脸风格与实时渲染。
- 自建多视角语音视频数据集,解决训练数据不足问题。
我们提出GaussianSpeech,一种从语音合成高保真、个性化的3D真人头部动画的新方法。为捕捉人类头部的丰富表情和细节(如皮肤褶皱、微小面部动作),将语音信号与3D高斯溅射相结合,生成时序连贯的运动序列。提出一种紧凑高效的3DGS头部表示,能生成依赖表情的颜色,并采用皱纹感知损失和人眼感知损失,以还原表情相关的皱纹等细节。为实现音频驱动的3D高斯点序列建模,设计了端到端的音频条件变压器模型,可直接从音频提取唇部与表情特征。由于缺乏高质量对齐的说话人音视频数据,我们采集了一个大规模多视角音频-视觉数据集,涵盖具有自然英语口音及多样面部结构的人群。GaussianSpeech在保持实时渲染速度的同时,持续达到当前最佳效果,生成视觉自然、表达丰富的动画。
原文摘要 · Abstract (English)
We introduce GaussianSpeech, a novel approach that synthesizes high-fidelity animation sequences of photo-realistic, personalized 3D human head avatars from spoken audio. To capture the expressive, detailed nature of human heads, including skin furrowing and finer-scale facial movements, we propose to couple speech signal with 3D Gaussian splatting to create realistic, temporally coherent motion sequences. We propose a compact and efficient 3DGS-based avatar representation that generates expression-dependent color and leverages wrinkle- and perceptually-based losses to synthesize facial details, including wrinkles that occur with different expressions. To enable sequence modeling of 3D Gaussian splats with audio, we devise an audio-conditioned transformer model capable of extracting lip and expression features directly from audio input. Due to the absence of high-quality datasets of talking humans in correspondence with audio, we captured a new large-scale multi-view dataset of audio-visual sequences of talking humans with native English accents and diverse facial geometry. GaussianSpeech consistently achieves state-of-the-art performance with visually natural motion at real time rendering rates, while encompassing diverse facial expressions and styles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。