arXiv:2507.09342cs.SDeess.AS2025-07

构建首个英-约鲁巴直接语音翻译语料库,助力低资源语言技术发展。

BENYO-S2ST-Corpus-1: A Bilingual English-to-Yoruba Direct Speech-to-Speech Translation Corpus

  • 用混合架构+AI合成音频,低成本构建英-约鲁巴语音对齐数据集。
  • 语料库含12,032对语音样本,总时长41.2小时,规模显著。
  • 可支持语音翻译、语音合成模型训练,适合非洲低资源语言研究者。

高资源语种到低资源语种(如英语到约鲁巴语)的语音到语音翻译(S2ST)数据集严重匮乏。为此,本研究构建了首个双语英语-约鲁巴直接语音翻译语料库BENYO-S2ST-Corpus-1。该语料库基于自研混合架构,实现大规模、低成本的S2ST语料生成。研究利用小规模(1,504样本)的YORULECT语料库中标准约鲁巴(SY)语音与转录文本,以及对应的英文(SE)转录文本,通过预训练模型Facebook MMS生成匹配的英文语音,并开发名为AcoustAug的音频增强算法,基于三个潜在声学特征生成增强语音。最终语料库包含每种语言12,032个音频样本,总计24,064个样本,总音频时长41.20小时。该语料库不仅可用于训练S2ST模型,还可用于预训练或改进现有模型。作为概念验证,研究基于该语料库与Coqui框架构建了约鲁巴语音合成模型YoruTTS-1.5,经1,000轮训练后基频均方根误差(F0 RMSE)为63.54,表明其与真实语音的基频相似度中等。BENYO-S2ST-Corpus-1与YoruTTS-1.5已公开发布(https://bit.ly/40bGMwi),可供研究人员用于多语言非洲低资源语言数据集构建,缩小高/低资源语言间的数字鸿沟。

原文摘要 · Abstract (English)

There is a major shortage of Speech-to-Speech Translation (S2ST) datasets for high resource-to-low resource language pairs such as English-to-Yoruba. Thus, in this study, we curated the Bilingual English-to-Yoruba Speech-to-Speech Translation Corpus Version 1 (BENYO-S2ST-Corpus-1). The corpus is based on a hybrid architecture we developed for large-scale direct S2ST corpus creation at reduced cost. To achieve this, we leveraged non speech-to-speech Standard Yoruba (SY) real-time audios and transcripts in the YORULECT Corpus as well as the corresponding Standard English (SE) transcripts. YORULECT Corpus is small scale(1,504) samples, and it does not have paired English audios. Therefore, we generated the SE audios using pre-trained AI models (i.e. Facebook MMS). We also developed an audio augmentation algorithm named AcoustAug based on three latent acoustic features to generate augmented audios from the raw audios of the two languages. BENYO-S2ST-Corpus-1 has 12,032 audio samples per language, which gives a total of 24,064 sample size. The total audio duration for the two languages is 41.20 hours. This size is quite significant. Beyond building S2ST models, BENYO-S2ST-Corpus-1 can be used to build pretrained models or improve existing ones. The created corpus and Coqui framework were used to build a pretrained Yoruba TTS model (named YoruTTS-1.5) as a proof of concept. The YoruTTS-1.5 gave a F0 RMSE value of 63.54 after 1,000 epochs, which indicates moderate fundamental pitch similarity with the reference real-time audio. Ultimately, the corpus architecture in this study can be leveraged by researchers and developers to curate datasets for multilingual high-resource-to-low-resource African languages. This will bridge the huge digital divides in translations among high and low-resource language pairs. BENYO-S2ST-Corpus-1 and YoruTTS-1.5 are publicly available at (https://bit.ly/40bGMwi).

语音翻译低资源语言数据集构建约鲁巴语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。