用少量语音样本生成藏语三种方言的自然语音,支持多说话人风格保留。
FMSD-TTS: Few-shot Multi-Speaker Multi-Dialect Text-to-Speech Synthesis for Ü-Tsang, Amdo and Kham Speech Dataset Generation
- 通过融合说话人与方言特征的模块,实现少样本下多方言语音合成。
- 在方言表现力和说话人相似度上显著优于基线模型。
- 开源了合成语料库与评估工具,助力藏语语音研究。
藏语是一种低资源语言,其三大方言——卫藏、安多和康巴——缺乏足够的平行语音语料,严重制约语音建模进展。为解决此问题,我们提出 FMSD-TTS,一种少样本、多说话人、多方言的文本转语音框架,仅需少量参考音频和显式方言标签即可生成平行方言语音。该方法引入创新的说话人-方言融合模块与方言专用动态路由网络(DSDR-Net),有效捕捉方言间的细微声学与语言差异,同时保持说话人身份一致性。大量客观与主观评估表明,FMSD-TTS 在方言表现力和说话人相似度上显著优于基线模型。我们进一步通过挑战性的语音到语音方言转换任务验证了合成语音的质量与实用性。本工作贡献包括:(1) 面向藏语多方言语音合成的新型少样本 TTS 系统;(2) 由 FMSD-TTS 生成的大规模合成藏语语音语料库公开发布;(3) 一个开源的标准化评估工具包,用于统一评估说话人相似性、方言一致性与音频质量。
原文摘要 · Abstract (English)
Tibetan is a low-resource language with minimal parallel speech corpora spanning its three major dialects-Ü-Tsang, Amdo, and Kham-limiting progress in speech modeling. To address this issue, we propose FMSD-TTS, a few-shot, multi-speaker, multi-dialect text-to-speech framework that synthesizes parallel dialectal speech from limited reference audio and explicit dialect labels. Our method features a novel speaker-dialect fusion module and a Dialect-Specialized Dynamic Routing Network (DSDR-Net) to capture fine-grained acoustic and linguistic variations across dialects while preserving speaker identity. Extensive objective and subjective evaluations demonstrate that FMSD-TTS significantly outperforms baselines in both dialectal expressiveness and speaker similarity. We further validate the quality and utility of the synthesized speech through a challenging speech-to-speech dialect conversion task. Our contributions include: (1) a novel few-shot TTS system tailored for Tibetan multi-dialect speech synthesis, (2) the public release of a large-scale synthetic Tibetan speech corpus generated by FMSD-TTS, and (3) an open-source evaluation toolkit for standardized assessment of speaker similarity, dialect consistency, and audio quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。