从单视角视频重建可动态变形、可真实光照的真人虚拟形象
Relightable and Dynamic Gaussian Avatar Reconstruction from Monocular Video
- 基于3D高斯溅射,引入动态皮肤权重模拟身体动作带来的形变
- 在稀疏视觉条件下仍能还原衣物褶皱等精细几何细节
- 支持任意光照下的真实渲染,适合虚拟人、影视特效场景
从单视角视频建模可重光照且可动画化的真人虚拟形象是一项长期且具有挑战性的任务。近期,神经辐射场(NeRF)和3D高斯溅射(3DGS)方法被用于该任务,但常因缺乏与身体运动相关的几何细节(如衣物褶皱)而产生不理想的逼真效果。本文提出一种基于3DGS的人体虚拟形象建模框架——可重光照动态高斯虚拟形象(RnD-Avatar),实现高保真度的姿态相关形变。为此,我们引入动态皮肤权重,根据姿态定义人体关节运动,并学习由身体动作引发的额外形变;同时提出一种新正则化项,在稀疏视觉线索下捕捉精细几何细节。此外,我们构建了一个包含多种光照条件的新多视角数据集以评估重光照能力。该框架可在新姿态和新视角下实现逼真渲染,并支持任意光照条件下的照片级照明效果。实验表明,该方法在新视角合成、新姿态渲染和重光照任务中均达到当前最优性能。
原文摘要 · Abstract (English)
Modeling relightable and animatable human avatars from monocular video is a long-standing and challenging task. Recently, Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS) methods have been employed to reconstruct the avatars. However, they often produce unsatisfactory photo-realistic results because of insufficient geometrical details related to body motion, such as clothing wrinkles. In this paper, we propose a 3DGS-based human avatar modeling framework, termed as Relightable and Dynamic Gaussian Avatar (RnD-Avatar), that presents accurate pose-variant deformation for high-fidelity geometrical details. To achieve this, we introduce dynamic skinning weights that define the human avatar's articulation based on pose while also learning additional deformations induced by body motion. We also introduce a novel regularization to capture fine geometric details under sparse visual cues. Furthermore, we present a new multi-view dataset with varied lighting conditions to evaluate relight. Our framework enables realistic rendering of novel poses and views while supporting photo-realistic lighting effects under arbitrary lighting conditions. Our method achieves state-of-the-art performance in novel view synthesis, novel pose rendering, and relighting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。