arXiv:2412.20148cs.CVcs.HC2024-12中稿 · ICASSP 2025被引 9

用分解式高斯场实现长发人物说话脸视频的逼真合成

DEGSTalk: Decomposed Per-Embedding Gaussian Fields for Hair-Preserving Talking Face Synthesis

  • 将高斯点分解为可动态调整的嵌入场,精准捕捉面部表情变化
  • 在公开数据集上生成视频的面部细节和长发运动更真实
  • 适合需要高保真长发人脸动画的研究者或开发者

准确合成说话人脸视频并保留长发等细微特征仍具挑战。为此,我们提出基于3D高斯溅射(3DGS)的分解式每嵌入高斯场(DEGSTalk),用于生成具有长发的真实感说话人脸视频。DEGSTalk采用可变形预嵌入高斯场,通过隐式表达系数动态调整预嵌入高斯原语,从而精确捕捉动态面部区域与细微表情。此外,提出动态长发保真肖像渲染技术,增强合成视频中长发运动的真实感。实验表明,相较于现有方法,DEGSTalk在复杂面部动态与长发保留方面均表现更优。代码将公开于 https://github.com/CVI-SZU/DEGSTalk。

原文摘要 · Abstract (English)

Accurately synthesizing talking face videos and capturing fine facial features for individuals with long hair presents a significant challenge. To tackle these challenges in existing methods, we propose a decomposed per-embedding Gaussian fields (DEGSTalk), a 3D Gaussian Splatting (3DGS)-based talking face synthesis method for generating realistic talking faces with long hairs. Our DEGSTalk employs Deformable Pre-Embedding Gaussian Fields, which dynamically adjust pre-embedding Gaussian primitives using implicit expression coefficients. This enables precise capture of dynamic facial regions and subtle expressions. Additionally, we propose a Dynamic Hair-Preserving Portrait Rendering technique to enhance the realism of long hair motions in the synthesized videos. Results show that DEGSTalk achieves improved realism and synthesis quality compared to existing approaches, particularly in handling complex facial dynamics and hair preservation. Our code will be publicly available at https://github.com/CVI-SZU/DEGSTalk.

人脸合成长发保留3D高斯

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。