用3D高斯点云实现实时音频驱动人脸生成,精度与速度兼备。
PGSTalker: Real-Time Audio-Driven Talking Head Generation via 3D Gaussian Splatting with Pixel-Aware Density Control
- 通过像素感知密度控制动态分配点云,提升关键区域细节
- 在公开数据集上实现毫秒级推理,唇音同步误差低于现有方法
- 适合虚拟主播、数字人等需要实时交互的场景
音频驱动的人脸生成在虚拟现实、数字人和影视制作中至关重要。尽管基于NeRF的方法能实现高保真重建,但渲染效率低且音画同步不佳。本文提出PGSTalker,一种基于3D高斯点云(3DGS)的实时音频驱动人脸合成框架。为提升渲染性能,提出像素感知密度控制策略,自适应分配点密度,在动态面部区域增强细节,同时减少冗余。此外,引入轻量级多模态门控融合模块,有效融合音频与空间特征,提高高斯点形变预测精度。在多个公开数据集上的实验表明,PGSTalker在渲染质量、唇音同步精度和推理速度方面均优于现有基于NeRF和3DGS的方法。该方法展现出强泛化能力,具备实际部署潜力。
原文摘要 · Abstract (English)
Audio-driven talking head generation is crucial for applications in virtual reality, digital avatars, and film production. While NeRF-based methods enable high-fidelity reconstruction, they suffer from low rendering efficiency and suboptimal audio-visual synchronization. This work presents PGSTalker, a real-time audio-driven talking head synthesis framework based on 3D Gaussian Splatting (3DGS). To improve rendering performance, we propose a pixel-aware density control strategy that adaptively allocates point density, enhancing detail in dynamic facial regions while reducing redundancy elsewhere. Additionally, we introduce a lightweight Multimodal Gated Fusion Module to effectively fuse audio and spatial features, thereby improving the accuracy of Gaussian deformation prediction. Extensive experiments on public datasets demonstrate that PGSTalker outperforms existing NeRF- and 3DGS-based approaches in rendering quality, lip-sync precision, and inference speed. Our method exhibits strong generalization capabilities and practical potential for real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。