arXiv:2509.11763cs.CV2025-09

通过多尺度融合与多属性学习,从单张随意拍摄照片重建高精度3D人脸。

MSMA: Multi-Scale Feature Fusion For Multi-Attribute 3D Face Reconstruction From Unconstrained Images

  • 设计多尺度特征融合框架,结合大核注意力机制提升特征提取精度。
  • 在MICC Florence、Facewarehouse等数据集上达到或超越当前最优结果。
  • 特别适合处理复杂表情、光照和姿态下的3D人脸重建任务。

从单张非约束图像中重建3D人脸仍具挑战性,主要源于环境条件多样。近年来,基于学习的方法通过捕捉复杂面部结构和细节取得显著进展,通常采用生成图像与输入图像间的投影损失来约束训练。然而,这些方法通常依赖大量标注的3D人脸数据,而这类数据获取困难且成本高昂。为减少对标注数据的依赖,许多方法使用投影损失进行训练。尽管如此,现有方法在多样面部属性和条件下仍难以捕捉细致的多尺度特征,导致重建不完整或不准确。本文提出一种多尺度特征融合与多属性学习相结合的MSMA框架,用于从非约束图像中进行3D人脸重建。该方法利用大核注意力模块增强跨尺度特征提取精度,实现从单张2D图像准确估计3D面部参数。在MICC Florence、Facewarehouse及自收集数据集上的综合实验表明,本方法性能与当前最先进水平相当,部分情况下甚至超越现有SOTA表现。

原文摘要 · Abstract (English)

Reconstructing 3D face from a single unconstrained image remains a challenging problem due to diverse conditions in unconstrained environments. Recently, learning-based methods have achieved notable results by effectively capturing complex facial structures and details across varying conditions. Consequently, many existing approaches employ projection-based losses between generated and input images to constrain model training. However, learning-based methods for 3D face reconstruction typically require substantial amounts of 3D facial data, which is difficult and costly to obtain. Consequently, to reduce reliance on labeled 3D face datasets, many existing approaches employ projection-based losses between generated and input images to constrain model training. Nonetheless, despite these advancements, existing approaches frequently struggle to capture detailed and multi-scale features under diverse facial attributes and conditions, leading to incomplete or less accurate reconstructions. In this paper, we propose a Multi-Scale Feature Fusion with Multi-Attribute (MSMA) framework for 3D face reconstruction from unconstrained images. Our method integrates multi-scale feature fusion with a focus on multi-attribute learning and leverages a large-kernel attention module to enhance the precision of feature extraction across scales, enabling accurate 3D facial parameter estimation from a single 2D image. Comprehensive experiments on the MICC Florence, Facewarehouse and custom-collect datasets demonstrate that our approach achieves results on par with current state-of-the-art methods, and in some instances, surpasses SOTA performance across challenging conditions.

3D人脸重建多尺度融合注意力机制单图重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。