用音素-口形对齐让人脸能说多国语言,效果更自然。
A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple Languages
- 用音素和口形作桥梁,跨语言生成更准的嘴型。
- 在12种语言上表现优秀,未见过的语言也能零样本适配。
- 构建了95小时多语言视频数据集,推动跨语言研究。
语音驱动的说话人脸生成(TFS)旨在根据语音输入生成逼真的面部动画。现有TFS模型在英语上表现良好,但在非英语语言中常出现嘴型不准、表情僵硬的问题,主要源于英语主导的训练数据集及缺乏跨语言泛化能力。为此,我们提出多语言专家(MuEx)框架,采用音素引导的专家混合(PG-MoE)架构,以音素和口形作为跨模态通用中间表示,缓解语言差异与数据偏见。进一步引入音素-口形对齐机制(PV-Align),强化音视频间的对应关系,提升同步性。同时构建包含12种语言、总计95.04小时高质量视频的多语言说话人脸数据集(MTFD),用于训练与评估。大量实验表明,MuEx在MTFD所有语言上均表现优异,并展现出对未见语言的零样本泛化能力。
原文摘要 · Abstract (English)
Speech-driven talking face synthesis (TFS) focuses on generating lifelike facial animations from speech input. Current TFS models perform well in English but struggle with non-English languages, producing inaccurate mouth shapes and rigid facial expressions. These limitations are mainly caused by English-dominated training datasets and the lack of cross-language generalization ability.To address these challenges, we propose Multilingual Experts (MuEx), a novel framework featuring a Phoneme-Guided Mixture-of-Experts (PG-MoE) architecture that employs phonemes and visemes as universal intermediaries to bridge the gap between audio and visual modalities, enabling lifelike multilingual TFS. We extract speech and visual features as phonemes and visemes, respectively, which represent the basic units of speech sounds and mouth movements, to alleviate linguistic differences and dataset bias.Furthermore, we introduce the Phoneme-Viseme Alignment Mechanism (PV-Align), which establishes robust cross-modal correspondences between phonemes and visemes to improve audiovisual synchronization. In addition, we construct a Multilingual Talking Face Dataset (MTFD) comprising 12 diverse languages with 95.04 hours of high-quality videos for training and evaluating multilingual TFS performance.Extensive experiments demonstrate that MuEx achieves superior performance across all languages in MTFD and exhibits effective zero-shot generalization to unseen languages without additional training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。