arXiv:2507.13875cs.CLeess.AS2025-07中稿 · Interspeech 2025

提升加泰罗尼亚语-西班牙语混用语音识别效果,用合成数据+语言标记优化模型。

Optimizing ASR for Catalan-Spanish Code-Switching: A Comparative Analysis of Methodologies

  • 通过生成合成混用数据、拼接单语音频、添加语言标记三种方法优化识别
  • 少量合成数据结合主语言标记使识别准确率最高,优于纯单语或真实混用数据
  • 成果适用于多语言社会中混用语境的语音识别,尤其适合媒体与政务场景

语言混用(CS)指交替使用两种或以上语言,因训练数据稀缺且语言相似性高,给自动语音识别(ASR)带来挑战。现有模型多依赖单语或混合语料库,难以反映真实混用模式,尤其在多语言社会中更为突出。以加泰罗尼亚语-西班牙语混用为例,该现象广泛存在于媒体和议会演讲中。本文探索三种改进策略:(1) 生成合成混用数据,(2) 拼接单语音频,(3) 利用真实混用数据并引入语言标记。从加泰罗尼亚语语料中提取混用数据,微调OpenAI Whisper模型,并发布于Hugging Face。结果表明,少量合成数据结合主导语言标记的方案表现最佳。

原文摘要 · Abstract (English)

Code-switching (CS), the alternating use of two or more languages, challenges automatic speech recognition (ASR) due to scarce training data and linguistic similarities. The lack of dedicated CS datasets limits ASR performance, as most models rely on monolingual or mixed-language corpora that fail to reflect real-world CS patterns. This issue is critical in multilingual societies where CS occurs in informal and formal settings. A key example is Catalan-Spanish CS, widely used in media and parliamentary speeches. In this work, we improve ASR for Catalan-Spanish CS by exploring three strategies: (1) generating synthetic CS data, (2) concatenating monolingual audio, and (3) leveraging real CS data with language tokens. We extract CS data from Catalan speech corpora and fine-tune OpenAI's Whisper models, making them available on Hugging Face. Results show that combining a modest amount of synthetic CS data with the dominant language token yields the best transcription performance.

语音识别混用语言Whisper多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。