用离散语音表征提升3D人脸动画精度,发现发音类别编码最有效。
From Tokens to Faces: Investigating Discrete Speech Representations for 3D Facial Animation

- 对比四类语音表征,发现发音类别编码更利于准确生成面部动作
- 在两种人脸解码器上均实现接近最优的面部重建质量
- 提出跨模态语音-面部生成流水线,共享离散表征空间
语音表征的选择对语音驱动的3D人脸动画至关重要。不同表征编码的信息不同:自监督学习(SSL)特征强调音段与语义线索,神经编解码器产生优化于声学重建的潜在表示,而自动语音识别(ASR)风格目标生成基于标签的空间。本文评估了四类语音表征在3D人脸合成中的表现,使用客观指标和感知评估,在两种人脸解码器上比较其面部重建质量。此外,还进行了探针分析,研究分词表示与音位单位及发音器官形变之间的关系。结果表明,编码发音类别有助于在语义和标签型表征上实现精准的人脸动画预测,且性能相当。基于此,我们引入一种视听文本转语音(AVTTS)流程,利用离散表示作为共享空间,同时解码语音与3D面部运动。
原文摘要 · Abstract (English)
The choice of speech representation is critical in speech-driven 3D facial animation. Representations differ in what they encode: SSL features emphasize segmental and semantic cues, neural codecs yield latents optimized for acoustic reconstruction, and ASR-style objectives produce label-based spaces. We evaluate four speech representation families for 3D facial synthesis, comparing their facial reconstruction quality across two facial decoders using objective metrics and a perceptual evaluation. We additionally conduct probing analyses that relate tokenized representations to phonetic units and to articulatory deformations. We found that encoding phonetic classes is beneficial for accurate facial animation prediction on both semantic and label-based representations with comparable facial animation quality. From the latter, we introduce an Audio Visual Text-to-Speech (AVTTS) pipeline that leverages, as a shared space, discrete representations to decode speech and 3D facial motion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。