用快速初始化+可学习修正图,分钟级生成高保真3D头像。
Gaussian Deja-vu: Creating Controllable 3D Gaussian Head-Avatars with Enhanced Generalization and Personalization Abilities
- 先用大规模2D数据训练通用头像模型,再用单视频微调个性化结果。
- 相比现有方法训练时间缩短至四分之一,生成速度提升至分钟级。
- 适合需要快速生成逼真3D头像的交互应用与内容创作场景。
3D高斯溅射(3DGS)在建模3D头像方面展现出巨大潜力,相较网格方法更具灵活性,比基于NeRF的方法更高效。然而,创建可控的3DGS头像仍耗时,常需数十分钟至数小时。为此,我们提出「Gaussian Deja-vu」框架:首先在大规模2D(合成与真实)图像数据集上训练通用头像模型,获得一个良好初始化的3D高斯头像;随后利用单目视频进行微调,实现个性化。个性化阶段引入可学习的表情感知修正混合图(expression-aware rectification blendmaps),无需神经网络即可快速收敛。实验表明,该方法在保真度上优于当前最优3DGS头像方法,训练时间减少至现有方法的四分之一,可在分钟内完成生成。
原文摘要 · Abstract (English)
Recent advancements in 3D Gaussian Splatting (3DGS) have unlocked significant potential for modeling 3D head avatars, providing greater flexibility than mesh-based methods and more efficient rendering compared to NeRF-based approaches. Despite these advancements, the creation of controllable 3DGS-based head avatars remains time-intensive, often requiring tens of minutes to hours. To expedite this process, we here introduce the "Gaussian Deja-vu" framework, which first obtains a generalized model of the head avatar and then personalizes the result. The generalized model is trained on large 2D (synthetic and real) image datasets. This model provides a well-initialized 3D Gaussian head that is further refined using a monocular video to achieve the personalized head avatar. For personalizing, we propose learnable expression-aware rectification blendmaps to correct the initial 3D Gaussians, ensuring rapid convergence without the reliance on neural networks. Experiments demonstrate that the proposed method meets its objectives. It outperforms state-of-the-art 3D Gaussian head avatars in terms of photorealistic quality as well as reduces training time consumption to at least a quarter of the existing methods, producing the avatar in minutes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。