arXiv:2509.14161cs.CLcs.SD2025-09被引 12

构建多语言混用语音数据集,助力低资源语言语音识别研究

CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset

  • 生成128小时跨语言混用语音训练数据,覆盖16组语言对
  • 包含4个测试集,共113种语言组合,覆盖高/低资源语言
  • 适合语音识别、机器翻译研究者,尤其关注多语混用场景

我们提出CS-FLEURS,一个用于开发和评估非高资源语言中混用语言语音识别与翻译系统的全新数据集。该数据集包含四个测试集,涵盖52种语言中的113种独特混用语言对:1)14组X-英语混用语言对,使用真实人声朗读合成生成的混用句子;2)16组X-英语混用语言对,采用生成式文本转语音技术;3)60组{阿拉伯语、中文、印地语、西班牙语}-X语言对,使用生成式文本转语音;4)45组低资源语言-英语混用语言对,采用拼接式文本转语音。此外,还提供128小时生成式文本转语音训练数据,覆盖16组X-英语语言对。我们希望该数据集能推动未来混用语言语音研究的扩展。数据集链接:https://huggingface.co/datasets/byan/cs-fleurs。

原文摘要 · Abstract (English)

We present CS-FLEURS, a new dataset for developing and evaluating code-switched speech recognition and translation systems beyond high-resourced languages. CS-FLEURS consists of 4 test sets which cover in total 113 unique code-switched language pairs across 52 languages: 1) a 14 X-English language pair set with real voices reading synthetically generated code-switched sentences, 2) a 16 X-English language pair set with generative text-to-speech 3) a 60 {Arabic, Mandarin, Hindi, Spanish}-X language pair set with the generative text-to-speech, and 4) a 45 X-English lower-resourced language pair test set with concatenative text-to-speech. Besides the four test sets, CS-FLEURS also provides a training set with 128 hours of generative text-to-speech data across 16 X-English language pairs. Our hope is that CS-FLEURS helps to broaden the scope of future code-switched speech research. Dataset link: https://huggingface.co/datasets/byan/cs-fleurs.

语音识别多语言代码混用数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。