首个大规模跨语言跨族裔混用数据集,助力多语言语音识别研究
SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset
- 构建多智能体框架合成真实混用语料,覆盖12种语言
- 包含42万条文本、80小时音频,来自63个族裔背景的174人
- 提出新评估指标SAER,更精准衡量混用场景性能
代码切换(CS)是对话或语句中交替使用两种及以上语言的现象,常受社会语境和说话者身份影响。这给通常针对单一语言设计的自动语音识别(ASR)系统带来挑战。随着多语言应用如代码切换语音识别(CSASR)、文本转语音(CSTTS)和跨语言信息检索(CLIR)需求增长,现有单语数据集已显不足。尽管已有部分代码切换数据集,但大多局限于同质族裔间的双语混合,缺乏大规模、多样化的基准。为此,我们提出LinguaMaster多智能体协作框架,用于高效可扩展的多语言数据合成。基于该框架,我们构建了首个大规模跨语言跨族裔代码切换数据集SwitchLingua:包含42万条跨语言文本样本,以及基于这些文本生成的超过80小时音频,来自174名来自18个国家/地区、63个种族/族裔背景的说话者。该数据集涵盖丰富的语言与文化多样性,为多语言与多元文化研究提供基础资源。此外,为解决现有ASR评估指标在代码切换场景下敏感性不足的问题,我们提出语义感知错误率(SAER),引入语义信息,实现更准确、上下文感知的系统性能评估。
原文摘要 · Abstract (English)
Code-switching (CS) is the alternating use of two or more languages within a conversation or utterance, often influenced by social context and speaker identity. This linguistic phenomenon poses challenges for Automatic Speech Recognition (ASR) systems, which are typically designed for a single language and struggle to handle multilingual inputs. The growing global demand for multilingual applications, including Code-Switching ASR (CSASR), Text-to-Speech (CSTTS), and Cross-Lingual Information Retrieval (CLIR), highlights the inadequacy of existing monolingual datasets. Although some code-switching datasets exist, most are limited to bilingual mixing within homogeneous ethnic groups, leaving a critical need for a large-scale, diverse benchmark akin to ImageNet in computer vision. To bridge this gap, we introduce \textbf{LinguaMaster}, a multi-agent collaboration framework specifically designed for efficient and scalable multilingual data synthesis. Leveraging this framework, we curate \textbf{SwitchLingua}, the first large-scale multilingual and multi-ethnic code-switching dataset, including: (1) 420K CS textual samples across 12 languages, and (2) over 80 hours of audio recordings from 174 speakers representing 18 countries/regions and 63 racial/ethnic backgrounds, based on the textual data. This dataset captures rich linguistic and cultural diversity, offering a foundational resource for advancing multilingual and multicultural research. Furthermore, to address the issue that existing ASR evaluation metrics lack sensitivity to code-switching scenarios, we propose the \textbf{Semantic-Aware Error Rate (SAER)}, a novel evaluation metric that incorporates semantic information, providing a more accurate and context-aware assessment of system performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。