用手机和手电筒在暗室拍脸,就能生成接近专业棚拍的皮肤材质图。
Facial Appearance Capture at Home with Patch-Level Reflectance Prior
- 在补丁级训练扩散模型,提升对低数据量高清扫描的泛化能力
- 通过匹配采集图像,生成无缝全分辨率面部反照率图,质量显著提升
- 适合想低成本实现高精度人脸数字化的普通用户
现有面部外观捕获方法可从智能手机录制的视频中重建合理的面部反照率,但质量仍远低于基于摄影棚拍摄的结果。本文提出一种日常可用的新方案,采用同位置手机与手电筒在昏暗环境中拍摄。为提升质量,关键观察是将面部反照率地图求解过程置于摄影棚扫描数据的分布内。具体地,我们首先在 Light Stage 扫描数据上学习一个扩散先验,再引导其生成最匹配捕获图像的反照率图。为改善泛化能力和训练稳定性,我们采用补丁级训练扩散先验,因为当前 Light Stage 数据集虽为超高清,但数据量有限。针对该先验,我们提出补丁级后验采样技术,从补丁级扩散模型中采样出无缝的全分辨率反照率图。实验表明,本方法大幅缩小了低成本拍摄与摄影棚录制之间的质量差距,使普通用户也能将自己数字化。代码将于 https://github.com/yxuhan/DoRA 发布。
原文摘要 · Abstract (English)
Existing facial appearance capture methods can reconstruct plausible facial reflectance from smartphone-recorded videos. However, the reconstruction quality is still far behind the ones based on studio recordings. This paper fills the gap by developing a novel daily-used solution with a co-located smartphone and flashlight video capture setting in a dim room. To enhance the quality, our key observation is to solve facial reflectance maps within the data distribution of studio-scanned ones. Specifically, we first learn a diffusion prior over the Light Stage scans and then steer it to produce the reflectance map that best matches the captured images. We propose to train the diffusion prior at the patch level to improve generalization ability and training stability, as current Light Stage datasets are in ultra-high resolution but limited in data size. Tailored to this prior, we propose a patch-level posterior sampling technique to sample seamless full-resolution reflectance maps from this patch-level diffusion model. Experiments demonstrate our method closes the quality gap between low-cost and studio recordings by a large margin, opening the door for everyday users to clone themselves to the digital world. Our code will be released at https://github.com/yxuhan/DoRA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。