通过语音同步技术提升AI配音的口型匹配度与语义准确性。
PS-TTS: Phonetic Synchronization in Text-to-Speech for Achieving Natural Automated Dubbing
- 用语言模型重写译文以匹配原声时长,实现时间对齐。
- 引入动态时间规整计算元音距离,使目标发音贴近源语音。
- 兼顾语义与发音相似性,适合多语言自动配音场景。
近年来,基于人工智能的配音技术取得进展,实现了将视频源语音自动转换为目标语言语音的自动化配音(AD)。然而,自然化配音仍面临时长与口型同步(lip-sync)等挑战,影响观感体验。为此,本文提出一种面向AD过程的文本重写同步方法,包含两步:等时性(isochrony)以满足时长约束,以及语音同步(PS)以保持口型一致。首先,利用语言模型重写译文,使目标语音时长与源语音一致;其次,提出PS方法,采用动态时间规整(DTW),并结合训练数据中元音距离的局部代价,使目标文本的发音与源语音元音更接近。进一步扩展为PSComet,联合考虑语义与语音相似性,更好地保留语义。所提方法集成于PS-TTS与PS-Comet TTS系统中。在韩语、英语唇读数据集及配音演员数据集上的评估表明,两者在多个客观指标上均优于无PS的TTS系统,并在韩-英、英-韩配音任务中超越人工配音员。扩展至法语后,测试所有语言对的跨语言适用性,结果显示:无论何种语言组合,PS-Comet均表现最优,在口型同步准确率与语义保留之间取得更好平衡,证明其比单独使用PS更具优势。
原文摘要 · Abstract (English)
Recently, artificial intelligence-based dubbing technology has advanced, enabling automated dubbing (AD) to convert the source speech of a video into target speech in different languages. However, natural AD still faces synchronization challenges such as duration and lip-synchronization (lip-sync), which are crucial for preserving the viewer experience. Therefore, this paper proposes a synchronization method for AD processes that paraphrases translated text, comprising two steps: isochrony for timing constraints and phonetic synchronization (PS) to preserve lip-sync. First, we achieve isochrony by paraphrasing the translated text with a language model, ensuring the target speech duration matches that of the source speech. Second, we introduce PS, which employs dynamic time warping (DTW) with local costs of vowel distances measured from training data so that the target text composes vowels with pronunciations similar to source vowels. Third, we extend this approach to PSComet, which jointly considers semantic and phonetic similarity to preserve meaning better. The proposed methods are incorporated into text-to-speech systems, PS-TTS and PS-Comet TTS. The performance evaluation using Korean and English lip-reading datasets and a voice-actor dubbing dataset demonstrates that both systems outperform TTS without PS on several objective metrics and outperform voice actors in Korean-to-English and English-to-Korean dubbing. We extend the experiments to French, testing all pairs among these languages to evaluate cross-linguistic applicability. Across all language pairs, PS-Comet performed best, balancing lip-sync accuracy with semantic preservation, confirming that PS-Comet achieves more accurate lip-sync with semantic preservation than PS alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。