用深度学习从多视角图像直接生成可动画的3D人脸网格
Multiple View Neural Regression of a Facial Shape Model
- 用合成多视角图像训练神经网络,直接预测重拓扑人脸网格
- 结合精确相机参数和3D关键点正则化,重建精度显著提升
- 适合影视动画、虚拟人领域需要高效建模的团队
生成可重拓扑的高质量3D人脸网格对高保真面部动画至关重要,但传统方法耗时费力。本文提出一种深度学习框架,直接从使用自研物理渲染系统Visage Craft生成的合成多视角图像中预测重拓扑人脸网格。该系统基于外观3D可变形模型(A3DMM),能输出即刻用于绑定与动画的标准网格,几乎无需人工干预。研究发现,引入准确的相机内参与外参可提升关键点定位精度与几何一致性;3D关键点正则化进一步优化了重建质量。所提方法有效降低了数据采集与处理成本,为高效生产级人脸建模提供了新路径。
原文摘要 · Abstract (English)
Creating re-topologized 3D facial meshes is essential for high-quality facial animation but remains labor-intensive and time-consuming. This dissertation explores more efficient approaches for capturing production-ready facial meshes through: (1) the development of VarIS, a custom light sphere for capturing high-resolution stereo geometry and reflectance maps; (2) analysis of camera parameters affecting automatic 2D and 3D landmarking; (3) synthetic-data methods for training neural face regression; and (4) techniques for improving neural multi-view face-shape regression. While VarIS enables photorealistic face capture, its operational and processing costs motivate a more scalable approach. A deep learning framework is therefore proposed to directly predict re-topologized facial meshes from synthetic multiview images generated with Visage Craft, an in-house physically based rendering system using an Appearance 3D Morphable Model (A3DMM). The system produces standardized meshes ready for rigging and animation with minimal human supervision. Results show that incorporating accurate camera intrinsics and extrinsics improves landmark accuracy and geometric consistency, while 3D landmark regularization further improves reconstruction quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。