arXiv:2506.22143cs.CLcs.SD2025-06中稿 · IEEE MLSP 2025被引 2

用拼接音频生成数据,提升低资源阿拉伯-英语混语语音识别性能

SAGE: Spliced-Audio Generated Data for Enhancing Foundational Models in Low-Resource Arabic-English Code-Switched Speech Recognition

  • 通过拼接音频合成人工混语数据,解决训练数据稀缺问题
  • 在混语基准上将词错误率降至31.1%,优于大型多语言模型
  • 适合低资源语言混合语音识别研究者和实际应用开发者

本文研究多种语音自监督学习模型在方言阿拉伯语及阿拉伯-英语混语语音上的表现。针对数据稀缺问题,提出改进的音频拼接方法生成人工混语数据。用该方法生成的SAGE数据微调已预训练模型,在阿拉伯-英语混语基准上使词错误率(WER)绝对降低7.8%。同时提出受经验回放启发的方法,提升模型在方言与混语间的泛化能力并缓解灾难性遗忘。引入域外3-gram语言模型后,平均总WER从31.7%降至26.6%。少样本微调进一步降低4.9%。最终在阿拉伯-英语混语基准上达到31.1%的WER,超越规模超十倍的USM和Whisper-large-v2模型,分别领先5.5%和8.4%。

原文摘要 · Abstract (English)

This paper investigates the performance of various speech SSL models on dialectal Arabic (DA) and Arabic-English code-switched (CS) speech. To address data scarcity, a modified audio-splicing approach is introduced to generate artificial CS speech data. Fine-tuning an already fine-tuned SSL model with the proposed Spliced-Audio Generated (SAGE) data results in an absolute improvement on Word Error Rate (WER) of 7.8% on Arabic and English CS benchmarks. Additionally, an Experience Replay (ER) inspired approach is proposed to enhance generalisation across DA and CS speech while mitigating catastrophic forgetting. Integrating an out-of-domain 3-gram language model reduces the overall mean WER from 31.7% to 26.6%. Few-shot fine-tuning for code-switching benchmarks further improves WER by 4.9%. A WER of 31.1% on Arabic-English CS benchmarks surpasses large-scale multilingual models, including USM and Whisper-large-v2 (both over ten times larger) by an absolute margin of 5.5% and 8.4%, respectively.

语音识别混语处理数据增强自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。