arXiv:2509.18060cs.CLcs.AI2025-09

构建统一藏语多方言语音合成框架,生成三大方言语音数据。

TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for Ü-Tsang, Amdo and Kham Speech Dataset Generation

  • 通过方言融合模块与动态路由网络捕捉多方言差异。
  • 在客观与主观评测中显著优于基线模型。
  • 可生成高质量语音,适用于方言转换任务。

藏语为低资源语言,其三大方言(卫藏、安多、康巴)缺乏平行语音语料库,制约了语音建模发展。为此,我们提出TMD-TTS,一种统一的藏语多方言文本转语音框架,能根据显式方言标签合成多方言语音。该方法包含方言融合模块与方言特化动态路由网络(DSDR-Net),有效捕捉方言间的声学与语言细微差异。大量客观与主观评估表明,TMD-TTS在方言表现力上显著优于基线模型。进一步通过具有挑战性的语音到语音方言转换(S2SDC)任务验证了合成语音的质量与实用性。

原文摘要 · Abstract (English)

Tibetan is a low-resource language with limited parallel speech corpora spanning its three major dialects (Ü-Tsang, Amdo, and Kham), limiting progress in speech modeling. To address this issue, we propose TMD-TTS, a unified Tibetan multi-dialect text-to-speech (TTS) framework that synthesizes parallel dialectal speech from explicit dialect labels. Our method features a dialect fusion module and a Dialect-Specialized Dynamic Routing Network (DSDR-Net) to capture fine-grained acoustic and linguistic variations across dialects. Extensive objective and subjective evaluations demonstrate that TMD-TTS significantly outperforms baselines in dialectal expressiveness. We further validate the quality and utility of the synthesized speech through a challenging Speech-to-Speech Dialect Conversion (S2SDC) task.

语音合成多方言低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。