用单相机实现高质量人脸捕获,速度达83帧/秒。
FaceSnap: Real-Time Personalized Lightstage Facial Performance Capture

- 分两阶段:先建个性化模型,再实时捕捉
- 83帧/秒生成4K动态纹理,细节更真实
- 适合影视动画快速迭代,无需多视角
Lightstage人脸捕获虽能生成高质量数字人,但耗时耗力。多相机阵列、数小时计算和海量数据存储成为流程瓶颈。本文提出FaceSnap,一种端到端框架,采用两阶段方法:首先通过一次运动范围序列的多视角优化,构建包含几何与表情相关外观的个性化模型;随后该模型支持仅用单个单目Lightstage相机实现实时高保真人脸性能捕获,无需再次多视角采集。FaceSnap以83帧/秒联合估计几何与动态4K纹理。4K纹理由一种新型个性化残差上采样器生成,可恢复个体特异性高频细节,通用上采样器无法捕捉。FaceSnap在几何精度上媲美全帧多视角优化,同时优于基于生产级3D数据训练的前馈方法,且仅依赖单摄像头视图。最后,我们发布了Multi4D,一个公开的光场环境4D人脸重建评估基准,支持跨方法拓扑不变的几何对比。
原文摘要 · Abstract (English)
Lightstage facial capture produces production-quality digital humans, but it is resource and labor-intensive. Multi-camera setups, hours of computation, and massive data storage create bottlenecks that hinder iterative workflows. This paper introduces FaceSnap, an end-to-end framework that streamlines capture via a two-stage approach. First, a one-time multi-view optimization from a range-of-motion sequence builds a personalized model encoding both geometry and expression-dependent appearance. This model then enables high-fidelity real-time facial performance capture from a single monocular lightstage camera, with no further multi-view capture required. FaceSnap jointly estimates geometry and dynamic 4K texture at 83 fps. The 4K texture is produced by a novel personalized residual upscaler that recovers subject-specific high-frequency detail, which generic upscalers fail to capture. FaceSnap achieves geometric accuracy competitive with full per-frame multi-view optimization while outperforming feed-forward methods trained on production-quality 3D data, all from a single camera view. Finally, we introduce Multi4D, a public benchmark for evaluating 4D facial reconstruction methods in lightstage environments, enabling topology-invariant geometric comparison across methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。