用合成数据训练多语混杂语音识别,无需真实混杂语料。
Can we train ASR systems on Code-switch without real code-switch data? Case study for Singapore's languages
- 通过短语级混合生成模拟自然混杂模式的合成数据。
- 在三种东南亚语言组合上,模型性能显著提升,马来-英语提升最明显。
- 为资源匮乏语言提供低成本构建混杂语音识别的新路径。
多语混杂(Code-switching, CS)在多语言场景中普遍存在,但因其语言复杂性导致标注数据稀少且昂贵,给语音识别(ASR)带来挑战。本研究探索仅使用合成混杂数据构建CS-ASR的方法。提出一种短语级混合策略,生成模仿自然混杂模式的合成数据,并利用单语数据与合成混杂数据联合微调大型预训练模型(Whisper、MMS、SeamlessM4T)。研究聚焦于三种资源匮乏的东南亚语言组合:马来-英语(BM-EN)、汉语-马来语(ZH-BM)、泰米尔-英语(TA-EN),建立了新的综合基准以评估主流ASR模型在混杂场景下的表现。实验结果表明,该训练策略在单语和混杂测试中均提升了性能,其中BM-EN提升最大,其次为TA-EN和ZH-BM。该方法为低成本开发CS-ASR提供了可行方案,对学术界与产业界均有价值。
原文摘要 · Abstract (English)
Code-switching (CS), common in multilingual settings, presents challenges for ASR due to scarce and costly transcribed data caused by linguistic complexity. This study investigates building CS-ASR using synthetic CS data. We propose a phrase-level mixing method to generate synthetic CS data that mimics natural patterns. Utilizing monolingual augmented with synthetic phrase-mixed CS data to fine-tune large pretrained ASR models (Whisper, MMS, SeamlessM4T). This paper focuses on three under-resourced Southeast Asian language pairs: Malay-English (BM-EN), Mandarin-Malay (ZH-BM), and Tamil-English (TA-EN), establishing a new comprehensive benchmark for CS-ASR to evaluate the performance of leading ASR models. Experimental results show that the proposed training strategy enhances ASR performance on monolingual and CS tests, with BM-EN showing highest gains, then TA-EN and ZH-BM. This finding offers a cost-effective approach for CS-ASR development, benefiting research and industry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。