arXiv:2409.11915eess.AS2024-09

用语调单元优化印地语语音合成,提升自然度与准确性。

Exploring an Inter-Pausal Unit (IPU) based Approach for Indic End-to-End TTS Systems

  • 以语调单元替代句子作为合成单元,适配印地语长句特点。
  • 相比传统方法,错误率降低,且生成语音更富韵律感。
  • 适合开发高自然度对话式语音系统,尤其适用于资源有限场景。

印度语言的句子通常比英语更长,且以语义完整短语为单位组合成句。长句导致文本到语音模型训练困难,合成语音韵律不佳。本文探索在端到端框架中采用语调单元(IPU)的方法,聚焦对话风格语音合成。研究对比了自回归的Tacotron2与非自回归的FastSpeech2架构,在印地语、泰米尔语和泰卢固语三种语言上进行实验。基于IPU的Tacotron2在合成音频中显著降低了插入与删除错误,为减少错误提供了新路径。该方法计算开销更小,生成语音韵律更丰富,优于传统的句子级系统。

原文摘要 · Abstract (English)

Sentences in Indian languages are generally longer than those in English. Indian languages are also considered to be phrase-based, wherein semantically complete phrases are concatenated to make up sentences. Long utterances lead to poor training of text-to-speech models and result in poor prosody during synthesis. In this work, we explore an inter-pausal unit (IPU) based approach in the end-to-end (E2E) framework, focusing on synthesising conversational-style text. We consider both autoregressive Tacotron2 and non-autoregressive FastSpeech2 architectures in our study and perform experiments with three Indian languages, namely, Hindi, Tamil and Telugu. With the IPU-based Tacotron2 approach, we see a reduction in insertion and deletion errors in the synthesised audio, providing an alternative approach to the FastSpeech(2) network in terms of error reduction. The IPU-based approach requires less computational resources and produces prosodically richer synthesis compared to conventional sentence-based systems.

语音合成端到端语调单元

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。