arXiv:2506.03884cs.CLcs.CV2025-06中稿 · INTERSPEECH 2025

利用语言亲缘关系,零样本合成印度小语种语音。

Kinship in Speech: Leveraging Linguistic Relatedness for Zero-Shot TTS in Indian Languages

  • 共享音素表征+文本解析规则适配目标语言音系
  • 在梵语、马拉地语等5种语言上生成可懂自然语音
  • 适合资源匮乏语言的快速语音合成,降低训练成本

文本到语音(TTS)系统通常需要高质量录音和精确转写。印度有1369种语言,其中22种为官方语言,使用13种文字。为所有语言训练TTS系统,尤其是缺乏数字资源的语言,任务艰巨。本文聚焦零样本语音合成,尤其针对文字与音系来自不同语系的语言。创新在于增强共享音素表征,并修改文本解析规则以匹配目标语言音系,从而降低合成器开销,实现快速适配。通过利用具有语言关联性的语种,成功生成了梵语、马哈拉施特拉语、卡纳拉孔卡尼语、迈蒂利语和库鲁克语的可懂且自然的语音。评估验证了该方法的有效性,展示了其在提升资源匮乏语言语音技术可及性方面的潜力。

原文摘要 · Abstract (English)

Text-to-speech (TTS) systems typically require high-quality studio data and accurate transcriptions for training. India has 1369 languages, with 22 official using 13 scripts. Training a TTS system for all these languages, most of which have no digital resources, seems a Herculean task. Our work focuses on zero-shot synthesis, particularly for languages whose scripts and phonotactics come from different families. The novelty of our work is in the augmentation of a shared phone representation and modifying the text parsing rules to match the phonotactics of the target language, thus reducing the synthesiser overhead and enabling rapid adaptation. Intelligible and natural speech was generated for Sanskrit, Maharashtrian and Canara Konkani, Maithili and Kurukh by leveraging linguistic connections across languages with suitable synthesisers. Evaluations confirm the effectiveness of this approach, highlighting its potential to expand speech technology access for under-represented languages.

语音合成零样本小语种语言亲缘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。