arXiv:2606.09553cs.CLcs.SD2026-06

构建37种低资源语言语音合成基准,推动公平的语音技术发展

OpenBibleTTS: Large-Scale Speech Resources and TTS Models for Low-Resource Languages

论文配图:OpenBibleTTS: Large-Scale Speech Resources and TTS Models for Low-Resource Languages
图 1 · 摘自论文原文
  • 基于圣经文本构建跨语言语音数据集,覆盖37种低资源语言
  • 多模型对比显示无单一系统在所有语言上表现最优,但通用模型更受听感青睐
  • 开源全部数据与模型,助力低资源语言语音研究

近年来神经文本转语音(TTS)和多语言语音生成取得显著进展,但这些成果在世界各语言间分布不均。现有模型仍以少数高资源语言为主,而多数低资源语言研究依赖人为降采样的高资源语料,无法反映真实低资源环境下拼写差异大、发音覆盖有限的问题。为此,我们提出OpenBibleTTS,一个涵盖37种被忽视语言的大规模低资源语音合成基准。在领域内圣经文本与领域外材料上,系统比较了多种TTS架构与大规模语音生成模型。结果表明,无单一系统在所有语言和指标上占优:Gemini-TTS在多数语言上获得最高听感评分,但仅在OpenBibleTTS上训练的单语言EveryVoice模型在可懂度方面依然最强,并在若干非洲语言中更受欢迎;而从零开始的开放模型在领域外文本上性能急剧下降,揭示出多语言覆盖与可靠合成质量之间的持续差距。研究结合自动评估与主观人类判断,并开源所有处理后的数据集、对齐信息及训练模型,以支持未来低资源语音合成研究。

原文摘要 · Abstract (English)

Recent advances in neural text-to-speech (TTS) and multilingual speech generation have substantially improved synthetic speech quality, yet these gains remain unevenly distributed across the world's languages. Existing models are still dominated by a small set of high-resource languages, while many studies of low-resource TTS are simulated on artificially downsampled high-resource corpora that do not reflect the orthographic variation and limited phonetic coverage encountered in genuinely underrepresented settings. As such, we introduce OpenBibleTTS, which is a large-scale benchmark for low-resource speech synthesis spanning 37 underrepresented languages. Moreover, a systematic comparison of various TTS architectures and large-scale speech generation models is conducted across in-domain Biblical text and out-of-domain material. Results show that no single system dominates across languages and metrics: Gemini-TTS achieves the highest listener ratings on most evaluated languages, but monolingual EveryVoice models trained on OpenBibleTTS remain strongest for intelligibility and are preferred in several African languages, while open from-scratch systems degrade sharply on out-of-domain text, revealing a persistent gap between broad multilingual coverage and reliable synthesis quality in underserved linguistic communities. We complement automatic evaluation with subjective human judgments, and open-source all processed datasets, alignments, and trained models to support future low-resource TTS research.

语音合成低资源语言数据集开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。