用视觉变压器从单图重建3D人脸,精度提升超15%。
Pixel3DMM: Versatile Screen-Space Priors for Single-Image 3D Face Reconstruction
- 基于DINO特征,预测每像素法向与纹理坐标作为几何约束
- 在超过97万张图像上训练,实现1000+身份的高精度重建
- 新基准首次评估姿态与中性表情,适合工业级人脸建模应用
本文解决从单张RGB图像重建3D人脸的问题。提出Pixel3DMM,一种基于视觉变压器的通用屏幕空间先验,通过预测每像素的几何线索来约束3D可变形模型(3DMM)的优化。利用DINO基础模型的潜在特征,引入专门设计的表面法向与uv坐标预测头。通过将三个高质量3D人脸数据集注册到FLAME网格拓扑,共获得超过1,000个身份和976,000张图像用于训练。针对3D人脸重建,提出一种基于FLAME的优化方法,从预测的uv坐标和法向估计中求解3DMM参数。为评估方法性能,构建了一个新基准,涵盖高多样性面部表情、视角与人种。关键的是,该基准首次同时评估有姿态与中性表情下的几何重建效果。最终,本方法在有姿态表情下几何精度超越最先进基线超过15%。
原文摘要 · Abstract (English)
We address the 3D reconstruction of human faces from a single RGB image. To this end, we propose Pixel3DMM, a set of highly-generalized vision transformers which predict per-pixel geometric cues in order to constrain the optimization of a 3D morphable face model (3DMM). We exploit the latent features of the DINO foundation model, and introduce a tailored surface normal and uv-coordinate prediction head. We train our model by registering three high-quality 3D face datasets against the FLAME mesh topology, which results in a total of over 1,000 identities and 976K images. For 3D face reconstruction, we propose a FLAME fitting opitmization that solves for the 3DMM parameters from the uv-coordinate and normal estimates. To evaluate our method, we introduce a new benchmark for single-image face reconstruction, which features high diversity facial expressions, viewing angles, and ethnicities. Crucially, our benchmark is the first to evaluate both posed and neutral facial geometry. Ultimately, our method outperforms the most competitive baselines by over 15% in terms of geometric accuracy for posed facial expressions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。