构建313小时真实场景多语种混用语音数据集,助力自然语言混用研究
CS-YODAS: A Mined Dataset of In-the-Wild Code-Switched Speech

- 基于YouTube数据,人工+算法协同挖掘真实混用语音
- 覆盖7种主语言,总量达313小时,含自然混用模式分析
- 适合研究多语种交流、语音识别与社会语言学的学者使用
我们提出CS-YODAS,一个基于创意共享许可的野外自然混用语音数据集,从多语种YouTube内容中挖掘而来。代码切换(CS)即在单次话语或对话中交替使用不同语言,常见于多语环境,但现有语音资源普遍规模小、领域单一或人为构造。基于YODAS语料库,我们开发了一套可扩展、人机协同的识别与验证流程,成功构建该数据集,总时长313小时,涵盖7种主语言,提供丰富真实的自发性混用语音实例。我们进一步分析了实际混用的分布特征,包括语言对频率与切换模式,并给出语音语言识别的基线结果。希望该数据集能推动更广泛深入的混用语音研究。数据链接:https://huggingface.co/datasets/byan/cs-yodas。
原文摘要 · Abstract (English)
We present CS-YODAS, a Creative Commons-licensed dataset of in-the-wild code-switched speech mined from multilingual YouTube data. Code-switching (CS), or the alternation between languages within an utterance or conversation, is common in multilingual settings but remains underrepresented in existing CS speech resources, which are typically small, domain-specific, or artificially constructed. Building on the YODAS corpus, we develop a scalable, human-in-the-loop pipeline for identifying and validating naturally occurring code-switching. The resulting dataset, which totals 313 hours and spans 7 matrix languages, provides diverse, real-world examples of spontaneous code-switched speech. We further analyze the distribution and characteristics of code-switching in the wild, examining language-pair frequencies and switching patterns, and report baseline results for spoken language identification. We hope that CS-YODAS will encourage broader and more comprehensive research on code-switched speech. Dataset link: https://huggingface.co/datasets/byan/cs-yodas.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。