单目图像生成逼真说话头像,无需额外训练即可适配新人脸。
Monocular and Generalizable Gaussian Talking Head Animation
- 利用深度信息和对称性补全单目图像的3D几何与外观特征。
- 两阶段预测策略提升不可见面部区域的参数精度,性能超越现有方法。
- 适合无多视角数据、需快速适配新人脸的实用场景。
本文提出Monocular and Generalizable Gaussian Talking Head Animation(MGGTalk),仅需单目数据即可实现说话头像动画生成,并能泛化至未见过的人脸身份而无需个性化微调。相比以往依赖多视角数据或繁琐微调的3D高斯泼溅(3DGS)方法,MGGTalk更具实用性与普适性。由于缺乏多视角与个性化训练数据,几何与外观信息不完整带来挑战。为此,MGGTalk引入深度信息以增强几何结构与面部对称性特征,通过像素级几何信息结合对称操作与点云过滤,确保3DGS位置参数的完整精确。随后采用两阶段策略,先预测可见面部区域的高斯参数,再利用其提升不可见区域的参数估计。大量实验表明,MGGTalk在多项指标上优于现有最优方法。
原文摘要 · Abstract (English)
In this work, we introduce Monocular and Generalizable Gaussian Talking Head Animation (MGGTalk), which requires monocular datasets and generalizes to unseen identities without personalized re-training. Compared with previous 3D Gaussian Splatting (3DGS) methods that requires elusive multi-view datasets or tedious personalized learning/inference, MGGtalk enables more practical and broader applications. However, in the absence of multi-view and personalized training data, the incompleteness of geometric and appearance information poses a significant challenge. To address these challenges, MGGTalk explores depth information to enhance geometric and facial symmetry characteristics to supplement both geometric and appearance features. Initially, based on the pixel-wise geometric information obtained from depth estimation, we incorporate symmetry operations and point cloud filtering techniques to ensure a complete and precise position parameter for 3DGS. Subsequently, we adopt a two-stage strategy with symmetric priors for predicting the remaining 3DGS parameters. We begin by predicting Gaussian parameters for the visible facial regions of the source image. These parameters are subsequently utilized to improve the prediction of Gaussian parameters for the non-visible regions. Extensive experiments demonstrate that MGGTalk surpasses previous state-of-the-art methods, achieving superior performance across various metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。