多语言语音驱动人脸动画,兼顾语种与个人风格
Polyglot: Multilingual Style Preserving Speech-Driven Facial Animation

- 用语音文本嵌入和参考表情序列提取风格特征
- 无需预设语言或说话人标签,支持跨语言跨人生成
- 联合建模语言与个性风格,动画更自然连贯
语音驱动人脸动画(SDFA)在影视、游戏和虚拟现实中有广泛应用。但现有模型大多基于单语数据训练,难以适应真实场景中的多语言需求。本文提出多语言个性化语音驱动动画框架Polyglot,通过语音文本嵌入编码语言信息,利用参考面部序列提取风格嵌入以捕捉个体说话特征。该方法不依赖预定义的语言或说话人标签,借助自监督学习实现跨语言、跨说话人的泛化能力。通过联合条件建模语言与风格,能有效捕捉节奏、发音方式及习惯性面部动作等表达特征,生成时序一致且逼真的动画。实验表明,该方法在单语和多语场景下均取得更好性能,为语音驱动动画中语言与个性风格的统一建模提供了新范式。
原文摘要 · Abstract (English)
Speech-Driven Facial Animation (SDFA) has gained significant attention due to its applications in movies, video games, and virtual reality. However, most existing models are trained on single-language data, limiting their effectiveness in real-world multilingual scenarios. In this work, we address multilingual SDFA, which is essential for realistic generation since language influences phonetics, rhythm, intonation, and facial expressions. Speaking style is also shaped by individual differences, not only by language. Existing methods typically rely on either language-specific or speaker-specific conditioning, but not both, limiting their ability to model their interaction. We introduce Polyglot, a unified diffusion-based architecture for personalized multilingual SDFA. Our method uses transcript embeddings to encode language information and style embeddings extracted from reference facial sequences to capture individual speaking characteristics. Polyglot does not require predefined language or speaker labels, enabling generalization across languages and speakers through self-supervised learning. By jointly conditioning on language and style, it captures expressive traits such as rhythm, articulation, and habitual facial movements, producing temporally coherent and realistic animations. Experiments show improved performance in both monolingual and multilingual settings, providing a unified framework for modeling language and personal style in SDFA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。