用1.5万小时弱标注语音训练出顶尖阿拉伯语语音识别模型
Advancing Arabic Speech Recognition Through Large-Scale Weakly Supervised Learning
- 基于Conformer架构,用1.5万小时弱标注数据训练
- 在标准评测中超越开源与闭源模型,达当前最佳水平
- 为低资源语言语音识别提供低成本可扩展方案
自动语音识别(ASR)在对话代理、工业机器人、呼叫中心自动化和自动字幕等场景中至关重要。然而,由于高质量标注语音数据稀缺且人工标注成本高昂,开发高性能的阿拉伯语ASR模型仍具挑战性。本文采用弱监督学习方法,基于Conformer架构,从零开始训练一个阿拉伯语ASR模型,使用涵盖现代标准阿拉伯语(MSA)和方言阿拉伯语(DA)的15,000小时弱标注语音数据,无需人工验证转录。尽管缺乏人工标注,该方法在标准基准测试中仍取得当前最优(SOTA)性能,优于现有开源与闭源模型。结果表明,弱监督学习是一种高效、可扩展的替代传统监督学习的方法,为低资源语言语音识别系统的发展提供了新路径。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) is crucial for human-machine interaction in diverse applications like conversational agents, industrial robotics, call center automation, and automated subtitling. However, developing high-performance ASR models remains challenging, particularly for low-resource languages like Arabic, due to the scarcity of large, labeled speech datasets, which are costly and labor-intensive to produce. In this work, we employ weakly supervised learning to train an Arabic ASR model using the Conformer architecture. Our model is trained from scratch on 15,000 hours of weakly annotated speech data covering both Modern Standard Arabic (MSA) and Dialectal Arabic (DA), eliminating the need for costly manual transcriptions. Despite the absence of human-verified labels, our approach achieves state-of-the-art (SOTA) results in Arabic ASR, surpassing both open and closed-source models on standard benchmarks. By demonstrating the effectiveness of weak supervision as a scalable, cost-efficient alternative to traditional supervised approaches, paving the way for improved ASR systems in low resource settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。