arXiv:2512.17773cs.CVcs.AI2025-12被引 1

用单图直接预测高保真人脸三维参数,速度快质量高

Pix2NPHM: Learning to Regress NPHM Reconstructions From a Single Image

  • 基于视觉变换器直接从单图回归NPHM参数
  • 在超过10万组3D数据上训练,支持实时重建
  • 适合需要高精度人脸重建的工业级应用

神经参数化头模型(NPHM)是替代基于网格的3D可变形模型的新进展,能实现更精细的几何细节。然而,由于其潜在空间表达能力强,将NPHM拟合到视觉输入极具挑战。为此,我们提出Pix2NPHM,一种基于视觉变换器(ViT)的网络,直接从单张图像回归NPHM参数。相比现有方法,该模型能重建更清晰的人脸几何结构和更准确的表情。为实现广泛泛化,我们采用在几何预测任务上预训练的领域专用ViT作为主干网络。模型在超过10万组NPHM注册数据(支持SDF空间直接监督)和大规模2D视频数据集(以法向估计作为伪真实几何)上进行训练。Pix2NPHM不仅可在交互帧率下实现3D重建,还可通过推理时优化表面法向与标准点图进一步提升几何保真度。最终,我们在真实场景数据上实现了前所未有的面部重建质量。

原文摘要 · Abstract (English)

Neural Parametric Head Models (NPHMs) are a recent advancement over mesh-based 3d morphable models (3DMMs) to facilitate high-fidelity geometric detail. However, fitting NPHMs to visual inputs is notoriously challenging due to the expressive nature of their underlying latent space. To this end, we propose Pix2NPHM, a vision transformer (ViT) network that directly regresses NPHM parameters, given a single image as input. Compared to existing approaches, the neural parametric space allows our method to reconstruct more recognizable facial geometry and accurate facial expressions. For broad generalization, we exploit domain-specific ViTs as backbones, which are pretrained on geometric prediction tasks. We train Pix2NPHM on a mixture of 3D data, including a total of over 100K NPHM registrations that enable direct supervision in SDF space, and large-scale 2D video datasets, for which normal estimates serve as pseudo ground truth geometry. Pix2NPHM not only allows for 3D reconstructions at interactive frame rates, it is also possible to improve geometric fidelity by a subsequent inference-time optimization against estimated surface normals and canonical point maps. As a result, we achieve unprecedented face reconstruction quality that can run at scale on in-the-wild data.

3D重建人脸建模视觉变换器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。