构建首个多说话人阿拉伯语语音合成数据集,支持语音合成与反欺诈研究。
ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis
- 整合真人录音、公开数据和商业合成语音,覆盖11位说话人。
- 总时长83.52小时,含7位真人说话人共约10小时语音。
- 适用于语音合成、语音克隆及深度伪造检测等任务。
我们提出ArVoice,一个包含带符号转录的多说话人现代标准阿拉伯语(MSA)语音语料库,旨在支持多说话人语音合成,并可应用于基于语音的符号恢复、语音转换和深度伪造检测等任务。ArVoice由三部分组成:(1) 六位具有不同人口统计特征的专业配音演员录制的新数据;(2) 经修改的阿拉伯语语音语料库子集;(3) 两个商用系统生成的高质量合成语音。整个语料库共包含83.52小时语音,涵盖11个说话人,其中约10小时为来自7位说话人的真人录音。我们训练了三个开源文本转语音(TTS)系统和两个语音转换系统,以展示该数据集的应用潜力。该语料库已向研究社区开放使用。
原文摘要 · Abstract (English)
We introduce ArVoice, a multi-speaker Modern Standard Arabic (MSA) speech corpus with diacritized transcriptions, intended for multi-speaker speech synthesis, and can be useful for other tasks such as speech-based diacritic restoration, voice conversion, and deepfake detection. ArVoice comprises: (1) a new professionally recorded set from six voice talents with diverse demographics, (2) a modified subset of the Arabic Speech Corpus; and (3) high-quality synthetic speech from two commercial systems. The complete corpus consists of a total of 83.52 hours of speech across 11 voices; around 10 hours consist of human voices from 7 speakers. We train three open-source TTS and two voice conversion systems to illustrate the use cases of the dataset. The corpus is available for research use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。