用4000小时自动标注数据训练阿拉伯语语音合成模型,无需依赖精准注音。
More Data, Fewer Diacritics: Scaling Arabic TTS
- 构建自动化流水线,收集并处理4000小时阿拉伯语语音数据。
- 4000小时非注音数据训练的模型性能接近注音数据训练结果。
- 适合关注低资源语言语音合成与自动化数据构建的研究者。
阿拉伯语文本转语音(TTS)研究受限于公开训练数据的匮乏及准确注音模型的缺失。本文通过探索在大规模自动标注数据上训练阿拉伯语TTS,构建了包含语音活动检测、语音识别、自动注音和噪声过滤的稳健数据处理流水线,最终获得约4000小时的阿拉伯语TTS训练数据。我们使用不同规模的数据(100、1000、4000小时)并对比有无注音的训练效果,发现尽管注音数据训练的模型整体表现更优,但大量非注音数据可显著弥补注音缺失带来的性能损失。我们计划发布一个无需注音即可运行的公开阿拉伯语TTS模型。
原文摘要 · Abstract (English)
Arabic Text-to-Speech (TTS) research has been hindered by the availability of both publicly available training data and accurate Arabic diacritization models. In this paper, we address the limitation by exploring Arabic TTS training on large automatically annotated data. Namely, we built a robust pipeline for collecting Arabic recordings and processing them automatically using voice activity detection, speech recognition, automatic diacritization, and noise filtering, resulting in around 4,000 hours of Arabic TTS training data. We then trained several robust TTS models with voice cloning using varying amounts of data, namely 100, 1,000, and 4,000 hours with and without diacritization. We show that though models trained on diacritized data are generally better, larger amounts of training data compensate for the lack of diacritics to a significant degree. We plan to release a public Arabic TTS model that works without the need for diacritization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。