用合成数据训练头部形象先验,少样本也能生成逼真新视角和表情。
Synthetic Prior for Few-Shot Drivable Head Avatar Inversion
- 基于大量合成头部数据构建先验模型,支持少样本微调。
- 在仅用几张图的情况下,生成效果超越现有单目与GAN方法。
- 适合需要隐私合规、少样本生成的虚拟人应用开发。
我们提出SynShot,一种基于合成先验的少样本可驱动头部形象逆向重建方法。针对三大挑战:1)可控3D生成网络训练需大量多样序列,但真实图像与高质量追踪网格难以获取;2)真实数据受隐私法规限制(如GDPR),需频繁删除模型与数据;3)现有单目头像模型泛化能力差,易过拟合特定视角分布。受仅用合成数据训练的机器学习模型启发,我们从大规模合成头部数据集(包含多样化身份、表情与视角)中学习先验模型。仅需少量输入图像,SynShot即可微调该先验以弥合域差距,建模出能泛化至新表情与新视角的逼真头部形象。采用3D高斯点阵表示,并结合卷积编码器-解码器输出UV纹理空间中的高斯参数。为应对头部不同区域建模复杂度差异(如皮肤与头发),引入显式控制机制,可按需增加各区域的原始数量。相比顶尖单目与GAN方法,SynShot在新视角与新表情合成上显著提升。
原文摘要 · Abstract (English)
We present SynShot, a novel method for the few-shot inversion of a drivable head avatar based on a synthetic prior. We tackle three major challenges. First, training a controllable 3D generative network requires a large number of diverse sequences, for which pairs of images and high-quality tracked meshes are not always available. Second, the use of real data is strictly regulated (e.g., under the General Data Protection Regulation, which mandates frequent deletion of models and data to accommodate a situation when a participant's consent is withdrawn). Synthetic data, free from these constraints, is an appealing alternative. Third, state-of-the-art monocular avatar models struggle to generalize to new views and expressions, lacking a strong prior and often overfitting to a specific viewpoint distribution. Inspired by machine learning models trained solely on synthetic data, we propose a method that learns a prior model from a large dataset of synthetic heads with diverse identities, expressions, and viewpoints. With few input images, SynShot fine-tunes the pretrained synthetic prior to bridge the domain gap, modeling a photorealistic head avatar that generalizes to novel expressions and viewpoints. We model the head avatar using 3D Gaussian splatting and a convolutional encoder-decoder that outputs Gaussian parameters in UV texture space. To account for the different modeling complexities over parts of the head (e.g., skin vs hair), we embed the prior with explicit control for upsampling the number of per-part primitives. Compared to SOTA monocular and GAN-based methods, SynShot significantly improves novel view and expression synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。