构建18.5小时中英混用语音数据集,助力双语语音识别研究
DOTA-ME-CS: Daily Oriented Text Audio-Mandarin English-Code Switching Dataset
- 基于34人录制数据,结合AI音色合成等技术增强多样性
- 含9300段音频,总计18.54小时,覆盖日常对话场景
- 适合语音识别、多语言处理领域研究者使用
代码切换(code-switching)指在交流中交替使用两种或以上语言,给自动语音识别(ASR)系统带来巨大挑战。现有模型与数据集在应对此类问题上能力有限。为填补这一空白并推动相关研究,我们发布DOTA-ME-CS:一个面向日常对话的中文-英文混用语音-文本数据集,包含18.54小时音频数据,共9,300段录音,来自34名参与者。为提升数据多样性,我们采用人工智能技术如音色合成、语速变化和噪声添加,增加任务复杂性与可扩展性。数据经精心筛选,兼顾多样性和质量,为研究双语语音识别提供可靠资源,并附详细数据分析。该数据集及配套代码将公开发布。
原文摘要 · Abstract (English)
Code-switching, the alternation between two or more languages within communication, poses great challenges for Automatic Speech Recognition (ASR) systems. Existing models and datasets are limited in their ability to effectively handle these challenges. To address this gap and foster progress in code-switching ASR research, we introduce the DOTA-ME-CS: Daily oriented text audio Mandarin-English code-switching dataset, which consists of 18.54 hours of audio data, including 9,300 recordings from 34 participants. To enhance the dataset's diversity, we apply artificial intelligence (AI) techniques such as AI timbre synthesis, speed variation, and noise addition, thereby increasing the complexity and scalability of the task. The dataset is carefully curated to ensure both diversity and quality, providing a robust resource for researchers addressing the intricacies of bilingual speech recognition with detailed data analysis. We further demonstrate the dataset's potential in future research. The DOTA-ME-CS dataset, along with accompanying code, will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。