arXiv:2510.13933eess.IV2025-10

用图像和法线图联合重建人脸驱动参数,精度高且泛化好。

Image-based Facial Rig Inversion

  • 双模态输入:颜色图与法线编码图分别用Hiera Transformer处理
  • 回归102个FACS驱动参数,合成与扫描数据均表现优异
  • 适合需要高保真人脸动画的数字人研发人员

我们提出一种基于图像的人脸驱动参数反演框架,融合RGB外观与RGB编码法线图两种模态。每种模态由独立的Hiera Transformer主干网络处理,提取特征后进行融合,用于回归102个源自面部动作编码系统(FACS)的驱动参数。在合成数据集和扫描数据集上的实验表明,该方法能良好泛化至扫描数据,实现高保真重建。

原文摘要 · Abstract (English)

We present an image-based rig inversion framework that leverages two modalities: RGB appearance and RGB-encoded normal maps. Each modality is processed by an independent Hiera transformer backbone, and the extracted features are fused to regress 102 rig parameters derived from the Facial Action Coding System (FACS). Experiments on synthetic and scanned datasets demonstrate that the method generalizes to scanned data, producing faithful reconstructions.

人脸重建FACS双模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。