arXiv:2601.18849cs.CV2026-01被引 1

用眨眼嵌入和哈希网格关键点编码,提升语音驱动人脸生成的嘴部精度。

Audio-Driven Talking Face Generation with Blink Embedding and Hash Grid Landmarks Encoding

  • 引入眨眼嵌入与哈希网格关键点编码,融合音频与面部特征
  • 在3D神经辐射场中实现更逼真的口型同步与动态表情
  • 适合需要高保真口型同步的虚拟人、数字主播场景

动态神经辐射场(NeRF)在生成高质量说话人3D肖像方面已取得显著进展。尽管渲染速度和生成质量大幅提升,但在准确高效地捕捉说话时的嘴部运动方面仍存在挑战。为此,本文提出一种基于眨眼嵌入与哈希网格关键点编码的自动方法,显著提升了说话人脸的保真度。具体而言,我们利用编码为条件特征的面部特征,并通过动态关键点变换器将音频特征作为残差项融入模型。同时,采用神经辐射场建模完整面部,实现逼真的面部表征。实验验证表明,该方法在生成质量上优于现有方法。

原文摘要 · Abstract (English)

Dynamic Neural Radiance Fields (NeRF) have demonstrated considerable success in generating high-fidelity 3D models of talking portraits. Despite significant advancements in the rendering speed and generation quality, challenges persist in accurately and efficiently capturing mouth movements in talking portraits. To tackle this challenge, we propose an automatic method based on blink embedding and hash grid landmarks encoding in this study, which can substantially enhance the fidelity of talking faces. Specifically, we leverage facial features encoded as conditional features and integrate audio features as residual terms into our model through a Dynamic Landmark Transformer. Furthermore, we employ neural radiance fields to model the entire face, resulting in a lifelike face representation. Experimental evaluations have validated the superiority of our approach to existing methods.

语音驱动人脸生成神经辐射场口型同步

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。