arXiv:2512.16511cs.CVcs.GR2025-12

提出多尺度注意力网络,实现人脸图像高精度光照分解与真实渲染。

Multi-scale Attention-Guided Intrinsic Decomposition and Rendering Pass Prediction for Facial Images

  • 采用分层残差编码与多尺度特征融合,提升反照率图边界清晰度。
  • 在512×512输入下生成1024×1024反照率图,六通道完整分解结果领先现有方法。
  • 适合需要高保真人脸重光照与材质编辑的应用场景。

在非约束光照条件下准确分解人脸图像的固有属性,是实现逼真重光照、高保真数字替身和增强现实效果的前提。本文提出MAGINet,一种多尺度注意力引导的固有属性网络,从单张RGB人像预测512×512光强归一化反照率图。MAGINet采用分层残差编码、瓶颈层中的空间-通道注意力以及解码器中的自适应多尺度特征融合,使反照率边界更锐利,光照不变性更强,优于以往U-Net变体。初始反照率图通过轻量三层CNN(RefinementNet)上采样至1024×1024并精修。基于该精修反照率,使用Pix2PixHD-based翻译器预测五项物理基础渲染通道:环境遮蔽、表面法线、镜面反射、半透明度及原始漫反射色(含残留光照)。结合精修反照率,共六通道构成完整固有分解。在FFHQ-UV-Intrinsics数据集上,联合使用掩码均方误差、VGG、边缘和块级LPIPS损失训练,全链路在反照率估计上达到当前最优性能,并显著提升完整渲染栈的保真度,支持高质量人脸重光照与材质编辑。

原文摘要 · Abstract (English)

Accurate intrinsic decomposition of face images under unconstrained lighting is a prerequisite for photorealistic relighting, high-fidelity digital doubles, and augmented-reality effects. This paper introduces MAGINet, a Multi-scale Attention-Guided Intrinsics Network that predicts a $512\times512$ light-normalized diffuse albedo map from a single RGB portrait. MAGINet employs hierarchical residual encoding, spatial-and-channel attention in a bottleneck, and adaptive multi-scale feature fusion in the decoder, yielding sharper albedo boundaries and stronger lighting invariance than prior U-Net variants. The initial albedo prediction is upsampled to $1024\times1024$ and refined by a lightweight three-layer CNN (RefinementNet). Conditioned on this refined albedo, a Pix2PixHD-based translator then predicts a comprehensive set of five additional physically based rendering passes: ambient occlusion, surface normal, specular reflectance, translucency, and raw diffuse colour (with residual lighting). Together with the refined albedo, these six passes form the complete intrinsic decomposition. Trained with a combination of masked-MSE, VGG, edge, and patch-LPIPS losses on the FFHQ-UV-Intrinsics dataset, the full pipeline achieves state-of-the-art performance for diffuse albedo estimation and demonstrates significantly improved fidelity for the complete rendering stack compared to prior methods. The resulting passes enable high-quality relighting and material editing of real faces.

图像分解人脸重光照扩散模型渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。