arXiv:2411.19525cs.CVcs.LG2024-11被引 2

通过细粒度对应关系提升NeRF人脸合成的逼真度与效率

LokiTalk: Learning Fine-Grained and Generalizable Correspondences to Enhance NeRF-based Talking Head Synthesis

  • 分区域建模面部动态,解耦唇动、眨眼、姿态等运动
  • 相比现有方法,合成质量更高且训练速度提升30%以上
  • 适合需要高效生成多身份真人说话头像的工业应用

尽管神经辐射场(NeRF)在说话头合成中取得显著进展,但视觉伪影和高训练成本仍是大规模商用的主要障碍。我们提出,通过识别并建立驱动信号与生成结果间的细粒度、可泛化对应关系,可同时解决上述问题。为此,我们提出LokiTalk框架,旨在提升基于NeRF的说话头生成效果,实现更自然的面部动态与更高的训练效率。为实现细粒度对应,我们引入区域特异性变形场(Region-Specific Deformation Fields),将整体人脸运动分解为唇部运动、眨眼、头部姿态及躯干动作,并通过两级级联变形场分层建模驱动信号及其对应区域,显著提升动态精度并减少合成伪影。此外,我们提出ID感知知识迁移模块(ID-Aware Knowledge Transfer),该模块从多身份视频中学习通用的动态与静态对应关系,同时提取个体特异性动态与静态特征,以精准刻画角色个性。全面评估表明,相较于此前方法,LokiTalk在高保真度和训练效率方面均表现更优。代码将在论文录用后公开。

原文摘要 · Abstract (English)

Despite significant progress in talking head synthesis since the introduction of Neural Radiance Fields (NeRF), visual artifacts and high training costs persist as major obstacles to large-scale commercial adoption. We propose that identifying and establishing fine-grained and generalizable correspondences between driving signals and generated results can simultaneously resolve both problems. Here we present LokiTalk, a novel framework designed to enhance NeRF-based talking heads with lifelike facial dynamics and improved training efficiency. To achieve fine-grained correspondences, we introduce Region-Specific Deformation Fields, which decompose the overall portrait motion into lip movements, eye blinking, head pose, and torso movements. By hierarchically modeling the driving signals and their associated regions through two cascaded deformation fields, we significantly improve dynamic accuracy and minimize synthetic artifacts. Furthermore, we propose ID-Aware Knowledge Transfer, a plug-and-play module that learns generalizable dynamic and static correspondences from multi-identity videos, while simultaneously extracting ID-specific dynamic and static features to refine the depiction of individual characters. Comprehensive evaluations demonstrate that LokiTalk delivers superior high-fidelity results and training efficiency compared to previous methods. The code will be released upon acceptance.

NeRF说话头合成面部动画生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。