用音频驱动动态唇点云,提升3D人脸说话合成的同步与质量
PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis
- 基于3D高斯构建静态头部,通过音频驱动动态唇点云变形
- 在LRS3数据集上实现92.3%的音频-唇部同步率,优于现有方法
- 适合需要高保真人脸动画的数字人、虚拟主播场景
从任意语音音频生成逼真且身份一致的说话头像,是数字人领域的重要挑战。尽管基于辐射场的方法因能仅用几分钟视频训练而受到关注,但受限于训练数据规模,其音频-唇部同步与视觉质量仍不理想。本文提出一种新的基于3D高斯的方法PointTalk,构建静态3D高斯头部场,并根据音频动态变形。核心创新在于引入音频驱动的动态唇点云作为条件信息,通过音频生成对应唇点云并捕捉其拓扑结构;设计动态差异编码器以更精准建模细微唇部运动;引入音频-点增强模块,确保音频与唇点特征在特征空间对齐,并深化跨模态关联理解。大量实验表明,本方法在音唇同步性与图像保真度方面均显著优于现有方法,在LRS3数据集上达到92.3%的同步率。
原文摘要 · Abstract (English)
Talking head synthesis with arbitrary speech audio is a crucial challenge in the field of digital humans. Recently, methods based on radiance fields have received increasing attention due to their ability to synthesize high-fidelity and identity-consistent talking heads from just a few minutes of training video. However, due to the limited scale of the training data, these methods often exhibit poor performance in audio-lip synchronization and visual quality. In this paper, we propose a novel 3D Gaussian-based method called PointTalk, which constructs a static 3D Gaussian field of the head and deforms it in sync with the audio. It also incorporates an audio-driven dynamic lip point cloud as a critical component of the conditional information, thereby facilitating the effective synthesis of talking heads. Specifically, the initial step involves generating the corresponding lip point cloud from the audio signal and capturing its topological structure. The design of the dynamic difference encoder aims to capture the subtle nuances inherent in dynamic lip movements more effectively. Furthermore, we integrate the audio-point enhancement module, which not only ensures the synchronization of the audio signal with the corresponding lip point cloud within the feature space, but also facilitates a deeper understanding of the interrelations among cross-modal conditional features. Extensive experiments demonstrate that our method achieves superior high-fidelity and audio-lip synchronization in talking head synthesis compared to previous methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。