用纹理空间高斯表示,实现4K级高保真动态人脸建模
TeGA: Texture Space Gaussian Avatars for High-Resolution Dynamic Head Modeling
- 在网格的连续UVD切空间中嵌入高斯点,实现精细外观建模
- 通过新型UVD变形场捕捉微表情与皱纹等高频细节,提升动态真实感
- 支持4K渲染,大幅增加3D高斯数量,适合虚拟人、元宇宙应用
基于3D高斯溅射的稀疏体素重建与渲染最近实现了可动画的高保真3D人脸虚拟人,可在任意视角下生成逼真的图像,成为远程通信、扩展现实和娱乐领域的重要技术。构建这类虚拟人需估计输入视频中面部各组件的复杂非刚性运动,但现有方法因运动估计不准确,导致可动画模型在保真度和细节上逊于单帧静态模型。此外,现有先进模型常受内存限制,使用较少3D高斯点,降低细节质量。为此,本文提出新型高细节3D人脸虚拟人模型,显著提升3D高斯数量并实现4K分辨率渲染。模型从多视角视频重建,基于网格化3D可变形模型提供粗粒度形变层;将3D高斯嵌入该网格的连续UVD切空间中,实现关键区域更有效的密集化。同时,通过新颖的UVD变形场对高斯进行形变,精准捕捉局部细微运动。核心贡献在于可变形高斯编码与整体拟合流程,使模型在保持外观细节的同时,有效捕获面部运动及皮肤皱纹等高频瞬时特征。
原文摘要 · Abstract (English)
Sparse volumetric reconstruction and rendering via 3D Gaussian splatting have recently enabled animatable 3D head avatars that are rendered under arbitrary viewpoints with impressive photorealism. Today, such photoreal avatars are seen as a key component in emerging applications in telepresence, extended reality, and entertainment. Building a photoreal avatar requires estimating the complex non-rigid motion of different facial components as seen in input video images; due to inaccurate motion estimation, animatable models typically present a loss of fidelity and detail when compared to their non-animatable counterparts, built from an individual facial expression. Also, recent state-of-the-art models are often affected by memory limitations that reduce the number of 3D Gaussians used for modeling, leading to lower detail and quality. To address these problems, we present a new high-detail 3D head avatar model that improves upon the state of the art, largely increasing the number of 3D Gaussians and modeling quality for rendering at 4K resolution. Our high-quality model is reconstructed from multiview input video and builds on top of a mesh-based 3D morphable model, which provides a coarse deformation layer for the head. Photoreal appearance is modelled by 3D Gaussians embedded within the continuous UVD tangent space of this mesh, allowing for more effective densification where most needed. Additionally, these Gaussians are warped by a novel UVD deformation field to capture subtle, localized motion. Our key contribution is the novel deformable Gaussian encoding and overall fitting procedure that allows our head model to preserve appearance detail, while capturing facial motion and other transient high-frequency features such as skin wrinkling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。