arXiv:2605.02496cs.SDcs.CL2026-05

用大模型解决藏语语音合成资源少、发音复杂难题

Tibetan-TTS:Low-Resource Tibetan Speech Synthesis with Large Model Adaptation

论文配图:Tibetan-TTS:Low-Resource Tibetan Speech Synthesis with Large Model Adaptation
图 1 · 摘自论文原文
  • 基于大模型+藏语文本适配+跨语言训练
  • 音质自然,发音准确率超96%,主观评分4.35
  • 适合低资源语言语音合成及多方言统一研究

藏语文语转换(TTS)长期受限于语音资源稀缺、方言差异大以及书面语与口语映射复杂等问题。本文首次在工业界提出基于大模型的藏语TTS系统,依托星晨AGI实验室开发的大规模语音合成模型,融合数据质量增强、面向藏语的文本表示与分词器适配、跨语言自适应训练等技术,实现低资源条件下的稳定生成。主观评测显示,音节级与BPE基系统主观评分(MOS)分别达4.28和4.35,发音准确率分别为97.6%和96.6%,优于外部商业藏语TTS接口。结果表明,结合大模型主干与藏语文本适配及跨语言训练,可实现高可用的低资源藏语语音合成,并为未来统一多方言藏语合成提供技术基础。

原文摘要 · Abstract (English)

Tibetan text-to-speech (TTS) has long been challenged by scarce speech resources, significant dialectal variation, and the complex mapping between written text and spoken pronunciation. To address these issues, this work presents, to the best of our knowledge, the first large-model-based Tibetan TTS system in the industry, built upon a large speech synthesis model developed by Xingchen AGI Lab. The proposed system integrates data quality enhancement, Tibetan-oriented text representation and tokenizer adaptation, and cross-lingual adaptive training for low-resource Tibetan speech synthesis. Experimental results show that the system can generate stable, natural, and intelligible Tibetan speech under low-resource conditions. In subjective evaluation, the MOS scores of the syllable-level and BPE-based systems reach 4.28 and 4.35, while their pronunciation accuracies reach 97.6% and 96.6%, respectively, outperforming an external commercial Tibetan TTS interface. These results demonstrate that combining a large-model backbone with Tibetan-oriented text representation adaptation and cross-lingual adaptive training enables highly usable low-resource Tibetan speech synthesis, and also provides a technical foundation for future unified multi-dialect Tibetan speech synthesis.

语音合成低资源大模型藏语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。