用声音单位建模面部动作,实现多语言自然口型同步
VQTalker: Towards Multilingual Talking Avatars through Facial Motion Tokenization
- 基于音素与视觉音素共性,用向量量化离散化面部运动
- 512×512分辨率下保持约11kbps低码率,生成质量领先
- 适合跨语言数字人、虚拟主播等需多语种口型同步的场景
我们提出VQTalker,一种基于向量量化(Vector Quantization)的多语言说话头生成框架,解决不同语言间唇同步与自然动作难题。该方法基于语音学原理:人类发音由有限音素构成,其对应的口部动作(可视音素)在多语言中具有共性。我们引入基于分组残差有限标量量化(GRFSQ)的面部运动分词器,将面部特征转化为离散表示,有效捕捉复杂动作并提升多语言泛化能力,即使训练数据有限。在此基础上,构建粗到细的运动生成流程,逐步细化面部动画。大量实验表明,VQTalker在视频驱动和语音驱动场景下均达当前最优性能,尤其在多语言设置中表现突出。方法可在512×512分辨率下实现约11kbps低码率,同时保持高质量输出。相关合成结果见https://x-lance.github.io/VQTalker。
原文摘要 · Abstract (English)
We present VQTalker, a Vector Quantization-based framework for multilingual talking head generation that addresses the challenges of lip synchronization and natural motion across diverse languages. Our approach is grounded in the phonetic principle that human speech comprises a finite set of distinct sound units (phonemes) and corresponding visual articulations (visemes), which often share commonalities across languages. We introduce a facial motion tokenizer based on Group Residual Finite Scalar Quantization (GRFSQ), which creates a discretized representation of facial features. This method enables comprehensive capture of facial movements while improving generalization to multiple languages, even with limited training data. Building on this quantized representation, we implement a coarse-to-fine motion generation process that progressively refines facial animations. Extensive experiments demonstrate that VQTalker achieves state-of-the-art performance in both video-driven and speech-driven scenarios, particularly in multilingual settings. Notably, our method achieves high-quality results at a resolution of 512*512 pixels while maintaining a lower bitrate of approximately 11 kbps. Our work opens new possibilities for cross-lingual talking face generation. Synthetic results can be viewed at https://x-lance.github.io/VQTalker.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。