用Transformer增强神经辐射场,让单目视频生成更逼真的动态人脸动画。
GAT-NeRF: Geometry-Aware-Transformer Enhanced Neural Radiance Fields for High-Fidelity 4D Facial Avatars
- 融合几何先验的轻量Transformer模块,提升面部细节建模能力。
- 在仅用单目视频条件下,显著改善皱纹和肤质等高频细节还原效果。
- 适合虚拟人、影视特效等领域,追求高保真动态人脸建模的研究者。
从单目视频高保真重建4D动态人脸是一项关键但极具挑战的任务,日益增长的沉浸式虚拟人应用需求推动其发展。尽管神经辐射场(NeRF)已推进场景表征,但在信息受限的单目流中捕捉动态皱纹、细微纹理等高频面部细节的能力仍需大幅提升。为此,我们提出一种新型混合神经辐射场框架——几何感知变压器增强型NeRF(GAT-NeRF),将Transformer机制融入NeRF流程。GAT-NeRF通过坐标对齐的多层感知机(MLP)与轻量级Transformer模块(称为几何感知变压器,GAT)协同工作,该模块融合3D空间坐标、3D可变形模型(3DMM)表达参数及可学习隐码等多模态输入,有效学习并增强与细粒度几何相关的特征表示。利用Transformer强大的特征学习能力,显著提升了对动态皱纹、痘疤等复杂局部面部模式的建模性能。全面实验表明,GAT-NeRF在视觉保真度和高频细节恢复方面均达到当前最优水平,为多媒体应用中构建真实动态数字人开辟新路径。
原文摘要 · Abstract (English)
High-fidelity 4D dynamic facial avatar reconstruction from monocular video is a critical yet challenging task, driven by increasing demands for immersive virtual human applications. While Neural Radiance Fields (NeRF) have advanced scene representation, their capacity to capture high-frequency facial details, such as dynamic wrinkles and subtle textures from information-constrained monocular streams, requires significant enhancement. To tackle this challenge, we propose a novel hybrid neural radiance field framework, called Geometry-Aware-Transformer Enhanced NeRF (GAT-NeRF) for high-fidelity and controllable 4D facial avatar reconstruction, which integrates the Transformer mechanism into the NeRF pipeline. GAT-NeRF synergistically combines a coordinate-aligned Multilayer Perceptron (MLP) with a lightweight Transformer module, termed as Geometry-Aware-Transformer (GAT) due to its processing of multi-modal inputs containing explicit geometric priors. The GAT module is enabled by fusing multi-modal input features, including 3D spatial coordinates, 3D Morphable Model (3DMM) expression parameters, and learnable latent codes to effectively learn and enhance feature representations pertinent to fine-grained geometry. The Transformer's effective feature learning capabilities are leveraged to significantly augment the modeling of complex local facial patterns like dynamic wrinkles and acne scars. Comprehensive experiments unequivocally demonstrate GAT-NeRF's state-of-the-art performance in visual fidelity and high-frequency detail recovery, forging new pathways for creating realistic dynamic digital humans for multimedia applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。