arXiv:2502.14178cs.GRcs.CV2025-02中稿 · ICASSP 2025被引 5

用3D先验解耦音频,实现自由视角的逼真人脸说话视频生成

NeRF-3DTalker: Neural Radiance Field with 3D Prior Aided Audio Disentanglement for Talking Head Synthesis

  • 引入3D先验信息,支持任意视角下的人脸视频生成
  • 通过音频解耦模块分离语音动作与说话风格,提升唇动同步精度
  • 设计局部全局归一化空间,解决生成帧位置偏移问题,适合高保真视频合成

说话头合成旨在根据音频生成唇同步的说话头视频。近期,NeRF在提升合成说话头的细节真实感方面表现出色,但现有基于音频的NeRF方法大多仅关注正面人脸渲染,难以生成新视角下的清晰视频。此外,当前3D说话头合成还面临声学与视觉空间对齐困难的问题,常导致唇动不同步。为此,我们提出NeRF-3DTalker:一种结合3D先验的音频解耦神经辐射场方法。该方法利用3D先验信息生成自由视角的清晰说话头;设计3D先验辅助音频解耦模块,将音频分为与3D语音动作相关的特征和与说话风格相关的特征;同时提出局部-全局标准化空间,从全局与局部语义层面归一化生成帧中偏离真实运动空间的位置。综合定性与定量实验表明,NeRF-3DTalker在生成逼真说话头视频方面优于当前最先进方法,图像质量与唇同步表现更优。

原文摘要 · Abstract (English)

Talking head synthesis is to synthesize a lip-synchronized talking head video using audio. Recently, the capability of NeRF to enhance the realism and texture details of synthesized talking heads has attracted the attention of researchers. However, most current NeRF methods based on audio are exclusively concerned with the rendering of frontal faces. These methods are unable to generate clear talking heads in novel views. Another prevalent challenge in current 3D talking head synthesis is the difficulty in aligning acoustic and visual spaces, which often results in suboptimal lip-syncing of the generated talking heads. To address these issues, we propose Neural Radiance Field with 3D Prior Aided Audio Disentanglement for Talking Head Synthesis (NeRF-3DTalker). Specifically, the proposed method employs 3D prior information to synthesize clear talking heads with free views. Additionally, we propose a 3D Prior Aided Audio Disentanglement module, which is designed to disentangle the audio into two distinct categories: features related to 3D awarded speech movements and features related to speaking style. Moreover, to reposition the generated frames that are distant from the speaker's motion space in the real space, we have devised a local-global Standardized Space. This method normalizes the irregular positions in the generated frames from both global and local semantic perspectives. Through comprehensive qualitative and quantitative experiments, it has been demonstrated that our NeRF-3DTalker outperforms state-of-the-art in synthesizing realistic talking head videos, exhibiting superior image quality and lip synchronization. Project page: https://nerf-3dtalker.github.io/NeRF-3Dtalker.

说话头生成NeRF音频解耦3D先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。