用高维高斯点提升人脸动画细节,让虚拟形象更逼真。
HyperGaussians: High-Dimensional Gaussian Splatting for High-Fidelity Animatable Face Avatars
- 将3D高斯点扩展为高维多变量高斯,增强表达能力。
- 在4个数据集19人上表现优于传统方法,细节更清晰。
- 适合做高保真人脸动画的开发者和虚拟现实研究者。
我们提出HyperGaussians,一种面向高质量可动画人脸头像的3D高斯点云新扩展。从单目视频生成精细人脸头像是增强现实与虚拟现实中的关键挑战,尽管静态人脸重建已取得显著进展,但现有方法在动态表情、复杂光照和微小细节上仍存在明显缺陷。当前主流方法3D高斯点云(3DGS)虽能高效渲染静态人脸,但在非线性形变、高频率细节如眼镜框、牙齿、反光等场景中表现不足。本文不依赖改进参数预测,而是重新思考高斯表示本身:提出高维多变量高斯,称为'HyperGaussians',通过可学习局部嵌入实现条件化,提升表达力。为解决高维协方差矩阵求逆带来的计算瓶颈,设计'逆协方差技巧'重参数化,显著提升效率。我们将HyperGaussians集成至当前最先进的快速单目人脸建模方法FlashAvatar,实验在4个数据集共19名受试者上验证,结果表明其在数值指标和视觉效果上均优于3DGS,尤其在高频率细节和复杂面部运动表现上优势明显。
原文摘要 · Abstract (English)
We introduce HyperGaussians, a novel extension of 3D Gaussian Splatting for high-quality animatable face avatars. Creating such detailed face avatars from videos is a challenging problem and has numerous applications in augmented and virtual reality. While tremendous successes have been achieved for static faces, animatable avatars from monocular videos still fall in the uncanny valley. The de facto standard, 3D Gaussian Splatting (3DGS), represents a face through a collection of 3D Gaussian primitives. 3DGS excels at rendering static faces, but the state-of-the-art still struggles with nonlinear deformations, complex lighting effects, and fine details. While most related works focus on predicting better Gaussian parameters from expression codes, we rethink the 3D Gaussian representation itself and how to make it more expressive. Our insights lead to a novel extension of 3D Gaussians to high-dimensional multivariate Gaussians, dubbed 'HyperGaussians'. The higher dimensionality increases expressivity through conditioning on a learnable local embedding. However, splatting HyperGaussians is computationally expensive because it requires inverting a high-dimensional covariance matrix. We solve this by reparameterizing the covariance matrix, dubbed the 'inverse covariance trick'. This trick boosts the efficiency so that HyperGaussians can be seamlessly integrated into existing models. To demonstrate this, we plug in HyperGaussians into the state-of-the-art in fast monocular face avatars: FlashAvatar. Our evaluation on 19 subjects from 4 face datasets shows that HyperGaussians outperform 3DGS numerically and visually, particularly for high-frequency details like eyeglass frames, teeth, complex facial movements, and specular reflections.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。