arXiv:2607.23657cs.CV2026-07

提出GRAPE模型,解决3D人脸网格估计中关节与表情混淆问题。

GRAPE: Graduated Routing for Articulated Portrait mesh Estimation

论文配图:GRAPE: Graduated Routing for Articulated Portrait mesh Estimation
图 1 · 摘自论文原文
  • 分阶段路由架构,融合躯干与头部运动先验
  • 显著提升姿态对齐与下颌-表情解耦精度
  • 适合虚拟形象生成与语音驱动口型动画

关节式人脸网格估计是3D理解、虚拟化身生成和沉浸式交互的基础。现有方法主要依赖3D可变形模型(3DMM),但以头为中心的模型受限于“悬浮头部”假设,将头部姿态误认为全局旋转,因缺乏颈部运动建模;而以躯干为中心的模型则难以表达高保真面部表情。此外,当前方法难以区分下颌运动与表情混合形状,常过度依赖表情来模拟张嘴。这些限制导致单目人脸重建在表示、监督和解剖参数估计上存在困难。为此,我们提出GRAPE(Graduated Routing for Articulated Portrait mesh Estimation)。构建包含显式躯干到头部运动链的肖像参数模型(PPM),并通过标准注入步骤融合FLAME与SMPL-X躯干。设计渐进式解剖对齐(PAA)网络,由预训练肖像编码器、分阶段掩码路由器及粗到精专家组成,遵循肖像解剖先验。采用多源监督训练:包括稀疏解剖关键点、特征蒸馏、前景掩码约束和相对几何约束。实验表明,GRAPE在网格恢复质量、姿态对齐和下颌-表情解耦方面优于现有方法。同时证明该方法可有效提升语音驱动说话头生成和3D肖像生成等下游任务性能。

原文摘要 · Abstract (English)

Articulated portrait mesh estimation is fundamental to 3D understanding, avatar generation, and immersive interaction. Existing approaches primarily rely on 3D Morphable Models (3DMMs). However, face-centric models suffer from the "floating head" assumption, conflating head pose with global rotation due to the lack of neck kinematics. Conversely, body-centric models lack high-fidelity facial expression capabilities. Furthermore, current methods struggle to disentangle jaw articulation from expression blendshapes, often over-relying on expressions for mouth opening. These limitations make monocular portrait recovery difficult across representation, supervision, and anatomical parameter estimation. To address these limitations, we introduce GRAPE(Graduated Routing for Articulated Portrait mesh Estimation). We build a Portrait Parametric Model (PPM) with an explicit torso-to-head kinematic chain and a canonical injection step to merge FLAME and the SMPL-X torso. We propose a Progressive Anatomical Alignment (PAA) network, which is composed of a pretrained portrait encoder, a Graduated-Mask Router, and coarse-to-fine experts that follow the portrait anatomical prior. We then train this network with multi-source supervision that combines sparse anatomical keypoints, feature distillation, foreground mask constraints, and relative geometry constraints. Experiments show that GRAPE improves portrait mesh recovery quality, pose alignment, and jaw--expression disentanglement over prior methods. We also demonstrate that our method can benefit the downstream tasks of audio-driven talking-head generation and 3D portrait generation.

3D人脸重建解剖建模表情解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。