用手机拍人脸也能高精度还原材质,靠的是新训练的光照先验模型。
Learning a Delighting Prior for Facial Appearance Capture in the Wild

- 用学习到的光照先验约束优化过程,摆脱对复杂光照的依赖。
- 在无标记视频上实现接近专业设备的材质还原效果,性能超越现有方法。
- 开源模型与4K可重光人脸数据集,助力数字人研究普及化。
高质量人脸外观捕获传统上依赖昂贵的影棚录制。近期工作采用手机拍摄的野外设置,但其基于模型的逆向渲染范式难以分离未知光照下的反射率。为此,我们提出将范式转向训练一个强大的愉悦性网络作为先验来约束优化。我们利用OLAT数据集和渲染的Light Stage扫描数据进行训练,提出数据集潜在调制(DLM)以无缝融合异构数据源。通过在核心网络中引入可学习的源感知令牌,我们解耦了数据集特有风格与物理愉悦性原理,从而涌现出超越现有专有模型的愉悦性先验。该先验使一个简单自动的外观捕获流程成为可能,仅需普通视频输入即可实现高质量反射率估计,显著优于先前方法。此外,我们利用此外观捕获方法将多视角NeRSemble数据集转化为NeRSemble-Scan,一个大规模的4K分辨率可重光人脸扫描集合。通过开源模型与NeRSemble-Scan数据集,我们推动高端人脸捕获的民主化,并为研究社区构建逼真数字人提供新基础。
原文摘要 · Abstract (English)
High-quality facial appearance capture has traditionally required costly studio recording. Recent works consider an in-the-wild smartphone-based setup; however, their model-based inverse rendering paradigm struggles with the complex disentanglement of reflectance from unknown illumination. To bridge this gap, we propose to shift the paradigm into training a powerful delighting network as a prior to constrain the optimization. We leverage the OLAT dataset and the rendered Light Stage scans for training, and propose Dataset Latent Modulation (DLM) to seamlessly integrate these heterogeneous data sources. Specifically, by conditioning the core network on learnable source-aware tokens, we decouple dataset-specific styles from physical delighting principles, enabling the emergence of a delighting prior that outperforms existing proprietary models. This powerful delighting prior enables a simple and automatic appearance capture pipeline that achieves high-quality reflectance estimation from casual video inputs, outperforming prior arts by a large margin. Furthermore, we leverage our appearance capture method to transform the multi-view NeRSemble dataset into NeRSemble-Scan, a large-scale collection of 4K-resolution relightable scans. By open-sourcing our model and the NeRSemble-Scan dataset, we democratize high-end facial capture and provide a new foundation for the research community to build photorealistic digital humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。