用多级注意力网络从单张照片重建高精度3D人脸。
Hierarchical MLANet: Multi-level Attention for 3D Face Reconstruction From Single Images
- 设计分层注意力机制,捕捉不同层级的面部特征。
- 在AFLW2000-3D和MICC Florence数据集上优于现有方法。
- 适合需要高效3D人脸重建的研究与应用者。
从野外拍摄的单张2D图像恢复3D人脸模型在计算机视觉领域受到广泛关注,因其具备广泛的应用潜力。然而,缺乏真实标注数据集以及真实环境的复杂性仍是主要挑战。本文提出一种基于卷积神经网络的层次化多级注意力网络(MLANet),实现从单张野外图像中重建3D人脸模型。该模型可预测精细的面部几何、纹理、姿态与光照参数。具体地,采用预训练的分层主干网络,并在2D人脸特征提取的不同阶段引入多级注意力机制。通过结合公开数据集中的3D形态模型(3DMM)参数与可微渲染器,采用半监督训练策略,实现端到端训练。在两个基准数据集AFLW2000-3D和MICC Florence上进行了大量实验,涵盖对比分析与消融研究,重点评估3D人脸重建与对齐任务的效果。结果表明,该方法在定量与定性评价上均表现优异。
原文摘要 · Abstract (English)
Recovering 3D face models from 2D in-the-wild images has gained considerable attention in the computer vision community due to its wide range of potential applications. However, the lack of ground-truth labeled datasets and the complexity of real-world environments remain significant challenges. In this chapter, we propose a convolutional neural network-based approach, the Hierarchical Multi-Level Attention Network (MLANet), for reconstructing 3D face models from single in-the-wild images. Our model predicts detailed facial geometry, texture, pose, and illumination parameters from a single image. Specifically, we employ a pre-trained hierarchical backbone network and introduce multi-level attention mechanisms at different stages of 2D face image feature extraction. A semi-supervised training strategy is employed, incorporating 3D Morphable Model (3DMM) parameters from publicly available datasets along with a differentiable renderer, enabling an end-to-end training process. Extensive experiments, including both comparative and ablation studies, were conducted on two benchmark datasets, AFLW2000-3D and MICC Florence, focusing on 3D face reconstruction and 3D face alignment tasks. The effectiveness of the proposed method was evaluated both quantitatively and qualitatively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。