arXiv:2412.19043cs.CLcs.AI2024-12中稿 · O-COCOSDA 2024被引 1

让语音合成同时说印尼语和英语混合句子,更自然清晰。

Indonesian-English Code-Switching Speech Synthesizer Utilizing Multilingual STEN-TTS and Bert LID

  • 用微调BERT识别每个词的语言,指导音素转换
  • 混合语句合成效果优于单一语言基线模型
  • 适合需要双语语音合成的印尼地区应用

多语言文本转语音系统可跨语言生成语音。在许多情况下,句子中会包含不同语言的片段,这种现象称为代码切换。这在印度尼西亚尤为常见,尤其是在印尼语和英语之间。尽管意义重大,但目前尚无研究开发出能够处理这两种语言间代码切换的多语言语音合成系统。本研究针对STEN-TTS中的印尼语-英语代码切换问题进行改进。关键修改包括:在文本到音素转换中加入语言识别组件,使用微调BERT实现逐词语言识别;以及从基础模型中移除语言嵌入。实验结果表明,该代码切换模型在自然度上表现更优,且语音可懂度相比印尼语和英语基线模型均有提升。

原文摘要 · Abstract (English)

Multilingual text-to-speech systems convert text into speech across multiple languages. In many cases, text sentences may contain segments in different languages, a phenomenon known as code-switching. This is particularly common in Indonesia, especially between Indonesian and English. Despite its significance, no research has yet developed a multilingual TTS system capable of handling code-switching between these two languages. This study addresses Indonesian-English code-switching in STEN-TTS. Key modifications include adding a language identification component to the text-to-phoneme conversion using finetuned BERT for per-word language identification, as well as removing language embedding from the base model. Experimental results demonstrate that the code-switching model achieves superior naturalness and improved speech intelligibility compared to the Indonesian and English baseline STEN-TTS models.

语音合成代码切换多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。