用风格化生成+高斯混合形状,实时打造多样卡通3D头像。
ToonifyGB: StyleGAN-based Gaussian Blendshapes for 3D Stylized Head Avatars
- 分两阶段:先用改进StyleGAN生成风格化视频,再构建风格化中性脸和表情混合形状。
- 在Arcane和Pixar两种风格下,生成的3D头像可任意切换表情,细节清晰流畅。
- 适合做动画角色生成、虚拟主播或游戏角色设计的开发者快速上手。
3D高斯混合形状的引入实现了从单目视频实时重建可驱动头像。Toonify是一种基于StyleGAN的面部图像风格化方法,应用广泛。为将Toonify扩展至利用高斯混合形状生成多样化风格化的3D头像,我们提出高效两阶段框架ToonifyGB。第一阶段(风格化视频生成)采用改进的StyleGAN,从输入视频帧生成风格化视频,克服了传统StyleGAN需固定分辨率裁剪对齐人脸的预处理限制,提升了风格化视频的稳定性,使高斯混合形状能更好捕捉视频帧的高频细节,为下一阶段高质量动画合成奠定基础。第二阶段(高斯混合形状合成)从生成的风格化视频中学习风格化中性头模型及一组表情混合形状。结合中性头模型与表情混合形状,ToonifyGB可高效渲染具有任意表情的风格化头像。我们在基准数据集上使用两种代表性风格(Arcane和Pixar)验证了ToonifyGB的有效性。
原文摘要 · Abstract (English)
The introduction of 3D Gaussian blendshapes has enabled the real-time reconstruction of animatable head avatars from monocular video. Toonify, a StyleGAN-based method, has become widely used for facial image stylization. To extend Toonify for synthesizing diverse stylized 3D head avatars using Gaussian blendshapes, we propose an efficient two-stage framework, ToonifyGB. In Stage 1 (stylized video generation), we adopt an improved StyleGAN to generate the stylized video from the input video frames, which overcomes the limitation of cropping aligned faces at a fixed resolution as preprocessing for normal StyleGAN. This process provides a more stable stylized video, which enables Gaussian blendshapes to better capture the high-frequency details of the video frames, facilitating the synthesis of high-quality animations in the next stage. In Stage 2 (Gaussian blendshapes synthesis), our method learns a stylized neutral head model and a set of expression blendshapes from the generated stylized video. By combining the neutral head model with expression blendshapes, ToonifyGB can efficiently render stylized avatars with arbitrary expressions. We validate the effectiveness of ToonifyGB on benchmark datasets using two representative styles: Arcane and Pixar.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。