arXiv:2409.11951cs.CVcs.GR2024-09被引 37

用分层建模实现高动态人脸实时渲染,支持复杂表情与大角度转动。

GaussianHeads: End-to-End Learning of Drivable Gaussian Head Avatars from Coarse-to-fine Representations

  • 分两阶段构建:先粗调模板网格,再在变形表面初始化并优化3D高斯点
  • 支持舌部形变、牙齿细节等精细结构在大幅运动下的高保真重建
  • 端到端训练可泛化至新表情和新视角,适合跨身份表情迁移应用

实时渲染人头虚拟形象是增强现实、视频游戏和电影等计算机图形应用的核心。现有方法虽能生成逼真图像,但在处理嘴内变化和大幅度头部姿态时表现不佳。本文提出一种新方法,从多视角图像中实时生成高度动态且可变形的人头虚拟形象。核心是分层次的头像表示,可捕捉面部表情和头部运动的复杂动态。首先从原始帧中提取丰富面部特征,对模板网格进行粗略变形;随后在变形表面上初始化3D高斯点,并在细粒度步骤中优化其位置。将粗到细的头像模型与可学习的头部姿态参数联合训练于端到端框架中。这不仅可通过视频输入实现可控的面部动画,还能在挑战性表情(如舌部形变和牙齿结构)下实现高保真新视角合成。此外,该方法在推理时具有良好泛化能力,适用于新表情和新姿态。我们在多个数据集上与相关方法对比验证了性能,涵盖不同身份的复杂表情序列。还展示了该方法在跨身份表情迁移中的应用潜力。

原文摘要 · Abstract (English)

Real-time rendering of human head avatars is a cornerstone of many computer graphics applications, such as augmented reality, video games, and films, to name a few. Recent approaches address this challenge with computationally efficient geometry primitives in a carefully calibrated multi-view setup. Albeit producing photorealistic head renderings, it often fails to represent complex motion changes such as the mouth interior and strongly varying head poses. We propose a new method to generate highly dynamic and deformable human head avatars from multi-view imagery in real-time. At the core of our method is a hierarchical representation of head models that allows to capture the complex dynamics of facial expressions and head movements. First, with rich facial features extracted from raw input frames, we learn to deform the coarse facial geometry of the template mesh. We then initialize 3D Gaussians on the deformed surface and refine their positions in a fine step. We train this coarse-to-fine facial avatar model along with the head pose as a learnable parameter in an end-to-end framework. This enables not only controllable facial animation via video inputs, but also high-fidelity novel view synthesis of challenging facial expressions, such as tongue deformations and fine-grained teeth structure under large motion changes. Moreover, it encourages the learned head avatar to generalize towards new facial expressions and head poses at inference time. We demonstrate the performance of our method with comparisons against the related methods on different datasets, spanning challenging facial expression sequences across multiple identities. We also show the potential application of our approach by demonstrating a cross-identity facial performance transfer application.

人脸建模3D高斯实时渲染表情动画

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。