arXiv:2409.12156cs.CV2024-09中稿 · BMVC 2024被引 3

用神经辐射场实现语音驱动的逼真人脸生成,表情与口型同步更自然。

JEAN: Joint Expression and Audio-guided NeRF-based Talking Face Generation

  • 基于NeRF架构,通过对比学习解耦语音与面部表情特征。
  • 在无真实标注数据下训练,仍能实现高保真口型同步和表情迁移。
  • 适合需要高质量语音驱动人脸动画的视频生成场景。

我们提出一种联合表情与语音引导的说话人脸生成新方法。现有方法或难以保持说话人身份,或无法生成真实表情。为此,我们设计基于NeRF的网络,在仅使用单目视频且无真实标签的情况下,学习音频与表情的解耦表征。首先通过自监督方式提取多说话人语音特征,并利用对比学习使音频特征对齐唇部运动,同时与面部其他肌肉运动解耦。随后设计基于Transformer的架构学习表情特征,捕捉长程面部动态并分离出与语音相关的嘴部动作。定量与定性评估表明,该方法可生成高保真说话人脸视频,在未见语音上实现了最先进的表情迁移与口型同步效果。

原文摘要 · Abstract (English)

We introduce a novel method for joint expression and audio-guided talking face generation. Recent approaches either struggle to preserve the speaker identity or fail to produce faithful facial expressions. To address these challenges, we propose a NeRF-based network. Since we train our network on monocular videos without any ground truth, it is essential to learn disentangled representations for audio and expression. We first learn audio features in a self-supervised manner, given utterances from multiple subjects. By incorporating a contrastive learning technique, we ensure that the learned audio features are aligned to the lip motion and disentangled from the muscle motion of the rest of the face. We then devise a transformer-based architecture that learns expression features, capturing long-range facial expressions and disentangling them from the speech-specific mouth movements. Through quantitative and qualitative evaluation, we demonstrate that our method can synthesize high-fidelity talking face videos, achieving state-of-the-art facial expression transfer along with lip synchronization to unseen audio.

说话人脸神经辐射场语音驱动表情同步

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。