用多视角法向预测实现秒级高保真人脸三维重建
Skullptor: High Fidelity 3D Head Reconstruction in Seconds with Multi-View Normal Prediction
- 通过跨视角注意力扩展单图模型,快速生成一致法向
- 结合逆渲染优化,恢复出接近密集拍摄的细节
- 仅需少量视角,计算量大幅降低,适合实时应用
从图像重建高保真3D人脸对众多应用至关重要,但现有方法存在根本局限。传统摄影测量虽细节出色,但需25-200+个视角、大量计算和人工清理(如面部毛发区域)。近期方法存在根本权衡:基础模型支持单图高效重建,但几何细节不足;优化方法虽更精细,却依赖密集视角和高算力。本文提出混合方法,融合两类范式优势。引入多视角表面法向预测模型,通过跨视角注意力机制,在前馈计算中生成几何一致的法向,扩展单图基础模型能力。随后将这些预测作为强几何先验,嵌入逆渲染优化框架,以恢复高频表面细节。该方法在单图与多视角对比中均优于现有技术,达到密集视点摄影测量水平的保真度,同时显著降低相机数量与计算成本。
原文摘要 · Abstract (English)
Reconstructing high-fidelity 3D head geometry from images is critical for a wide range of applications, yet existing methods face fundamental limitations. Traditional photogrammetry achieves exceptional detail but requires extensive camera arrays (25-200+ views), substantial computation, and manual cleanup in challenging areas like facial hair. Recent alternatives present a fundamental trade-off: foundation models enable efficient single-image reconstruction but lack fine geometric detail, while optimization-based methods achieve higher fidelity but require dense views and expensive computation. We bridge this gap with a hybrid approach that combines the strengths of both paradigms. Our method introduces a multi-view surface normal prediction model that extends monocular foundation models with cross-view attention to produce geometrically consistent normals in a feed-forward pass. We then leverage these predictions as strong geometric priors within an inverse rendering optimization framework to recover high-frequency surface details. Our approach outperforms state-of-the-art single-image and multi-view methods, achieving high-fidelity reconstruction on par with dense-view photogrammetry while reducing camera requirements and computational cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。