arXiv:2502.18913cs.CLcs.SD2025-02被引 12

104小时真实中英混用对话数据集,助力语音识别模型提升跨语言能力。

CS-Dialogue: A 104-Hour Dataset of Spontaneous Mandarin-English Code-Switching Dialogues for Speech Recognition

  • 收集200人自发性中英混用对话,提供完整录音与转写。
  • 104小时数据展现自然混用模式,现有模型仍存显著识别误差。
  • 适合语音识别、多语言研究者,推动真实场景下ASR技术发展。

语码转换(Code-switching, CS)指在单次对话中交替使用两种或以上语言,对自动语音识别(ASR)系统构成重大挑战。现有中英混用数据集普遍存在规模小、非自发、缺乏完整对话录音与转写等问题,制约了真实对话场景下鲁棒性ASR模型的发展。本文提出CS-Dialogue,一个大规模中英混用语音数据集,包含200名说话者产生的104小时自发对话。与以往数据集不同,该数据集提供完整的对话录音与转写,捕捉连续语音中的自然语码转换模式。我们详细描述了数据采集与标注流程,给出数据集统计信息,并使用Transformer、Conformer和Branchformer等先进模型建立基准测试性能。实验表明,语码转换语音识别仍具挑战性,现有预训练模型如Whisper仍有较大改进空间。该数据集将向学术界免费开放。

原文摘要 · Abstract (English)

Code-switching (CS), the alternation between two or more languages within a single conversation, presents significant challenges for automatic speech recognition (ASR) systems. Existing Mandarin-English code-switching datasets often suffer from limitations in size, spontaneity, and the lack of full-length dialogue recordings with transcriptions, hindering the development of robust ASR models for real-world conversational scenarios. This paper introduces CS-Dialogue, a novel large-scale Mandarin-English code-switching speech dataset comprising 104 hours of spontaneous conversations from 200 speakers. Unlike previous datasets, CS-Dialogue provides full-length dialogue recordings with complete transcriptions, capturing naturalistic code-switching patterns in continuous speech. We describe the data collection and annotation processes, present detailed statistics of the dataset, and establish benchmark ASR performance using state-of-the-art models. Our experiments, using Transformer, Conformer, and Branchformer, demonstrate the challenges of code-switching ASR, and show that existing pre-trained models such as Whisper still have the space to improve. The CS-Dialogue dataset will be made freely available for all academic purposes.

语音识别语码转换多语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。