用手机视频在野外捕捉高精度人脸反照率,效果逼近可控光照下表现。
WildCap: Facial Albedo Capture in the Wild via Hybrid Inverse Rendering
- 结合数据驱动与模型驱动方法,分离复杂光照下的人脸反照率。
- 在野外拍摄条件下,反照率还原质量显著优于现有方法。
- 适合移动端人脸建模、虚拟形象生成等实际应用。
现有方法在可控光照下可实现高质量人脸反照率捕获,但成本高且适用性受限。本文提出WildCap,一种从手机拍摄的野外视频中实现高精度人脸反照率捕获的新方法。为分离野外复杂光照下的高质量反照率,提出一种混合逆渲染框架:首先使用数据驱动方法SwitchLight将图像转换为更受控的条件,再采用基于模型的逆渲染。然而,网络预测中的局部伪影(如阴影烘焙)非物理,影响光照与材质的准确重建。为此,提出新型纹理网格光照模型,将非物理效应解释为干净反照率被局部物理光照照射的结果。优化过程中联合采样扩散先验以约束反照率图,并优化光照,有效解决局部光与反照率之间的尺度模糊问题。最终,其他反射率图由反照率推导得出。本方法在相同采集条件下显著优于以往方法,大幅缩小了野外与可控光照录制间的质量差距。
原文摘要 · Abstract (English)
Existing methods achieve high-quality facial albedo capture under controllable lighting, which increases capture cost and limits usability. We propose WildCap, a novel method for high-quality facial albedo capture from a smartphone video recorded in the wild. To disentangle high-quality albedo from complex lighting effects in in-the-wild captures, we propose a novel hybrid inverse rendering framework. We first apply a data-driven method, i.e., SwitchLight, to convert the captured images into more constrained conditions and then adopt model-based inverse rendering. However, unavoidable local artifacts in network predictions, such as shadow-baking, are non-physical and thus hinder accurate inverse rendering of lighting and material. To address this, we propose a novel texel grid lighting model to explain non-physical effects as clean albedo illuminated by local physical lighting. During optimization, we jointly sample a diffusion prior for the albedo map and optimize the lighting, effectively resolving scale ambiguity between local lights and albedo. Other reflectance maps are then predicted from the albedo. Our method achieves significantly better results than prior arts in the same capture setup, closing the quality gap between in-the-wild and controllable recordings by a large margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。