arXiv:2502.10982cs.CV2025-02ICLR被引 17

提升单图3D人脸表情重建精度,解决嘴型不规则等问题

TEASER: Token Enhanced Spatial Modeling for Expressions Reconstruction

  • 用多尺度令牌提取面部外观信息,结合神经渲染提供几何引导
  • 引入姿态相关关键点损失,显著改善细微表情定位准确性
  • 输出可解释的令牌,适合视频驱动、表情迁移等下游任务

从一张自然场景下的单图进行3D人脸重建是人本计算机视觉中的关键任务。现有方法虽能恢复准确的面部形状,但在精细表情捕捉方面仍有提升空间,尤其在嘴型不规则、表情夸张及面部不对称运动时表现不佳。本文提出TEASER(Token EnhAnced Spatial modeling for Expressions Reconstruction),针对现有方法在自重建中光度损失不足和细微表情定位不准两大问题,引入多尺度令牌化器提取面部外观信息,结合神经渲染生成精确的几何指导。此外,还设计了姿态依赖的关键点损失以进一步优化几何性能。实验结果表明,TEASER在多个数据集上均达到当前最优表达重建效果,且输出的可解释令牌适用于照片级真实感人脸视频驱动、表情迁移与身份替换等应用。

原文摘要 · Abstract (English)

3D facial reconstruction from a single in-the-wild image is a crucial task in human-centered computer vision tasks. While existing methods can recover accurate facial shapes, there remains significant space for improvement in fine-grained expression capture. Current approaches struggle with irregular mouth shapes, exaggerated expressions, and asymmetrical facial movements. We present TEASER (Token EnhAnced Spatial modeling for Expressions Reconstruction), which addresses these challenges and enhances 3D facial geometry performance. TEASER tackles two main limitations of existing methods: insufficient photometric loss for self-reconstruction and inaccurate localization of subtle expressions. We introduce a multi-scale tokenizer to extract facial appearance information. Combined with a neural renderer, these tokens provide precise geometric guidance for expression reconstruction. Furthermore, TEASER incorporates a pose-dependent landmark loss to further improve geometric performances. Our approach not only significantly enhances expression reconstruction quality but also offers interpretable tokens suitable for various downstream applications, such as photorealistic facial video driving, expression transfer, and identity swapping. Quantitative and qualitative experimental results across multiple datasets demonstrate that TEASER achieves state-of-the-art performance in precise expression reconstruction.

3D人脸重建表情捕捉神经渲染可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。