首次实现语音驱动的高保真动态纹理人脸动画
Towards High-fidelity 3D Talking Avatar with Personalized Dynamic Texture
- 用扩散模型同步生成面部动作与8K动态纹理
- 在100人、100分钟的4D数据集上验证效果
- 支持风格解耦控制,适合影视/虚拟主播场景
语音驱动的3D人脸动画取得显著进展,但多数方法仅关注网格/几何运动,忽视动态纹理的影响。本文揭示动态纹理在高保真说话头像渲染中的关键作用,构建了首个高分辨率4D数据集TexTalk4D,包含100名受试者共100分钟音频同步扫描级网格及8K动态纹理。基于该数据集,我们探索了运动与纹理间的内在关联,提出基于扩散的框架TexTalker,可从语音同步生成面部运动与动态纹理。进一步提出基于枢轴的风格注入策略,实现纹理与运动风格的解耦控制。TexTalker是首个实现音频同步面部运动与动态纹理生成的方法,在运动合成上优于现有技术,同时生成与面部动作一致的逼真纹理。项目页:https://xuanchenli.github.io/TexTalk/
原文摘要 · Abstract (English)
Significant progress has been made for speech-driven 3D face animation, but most works focus on learning the motion of mesh/geometry, ignoring the impact of dynamic texture. In this work, we reveal that dynamic texture plays a key role in rendering high-fidelity talking avatars, and introduce a high-resolution 4D dataset \textbf{TexTalk4D}, consisting of 100 minutes of audio-synced scan-level meshes with detailed 8K dynamic textures from 100 subjects. Based on the dataset, we explore the inherent correlation between motion and texture, and propose a diffusion-based framework \textbf{TexTalker} to simultaneously generate facial motions and dynamic textures from speech. Furthermore, we propose a novel pivot-based style injection strategy to capture the complicity of different texture and motion styles, which allows disentangled control. TexTalker, as the first method to generate audio-synced facial motion with dynamic texture, not only outperforms the prior arts in synthesising facial motions, but also produces realistic textures that are consistent with the underlying facial movements. Project page: https://xuanchenli.github.io/TexTalk/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。