arXiv:2508.19587cs.CLcs.AI2025-08被引 1

提升阿拉伯语发音评估的稳定性,解决孤立字母识别难题

Towards stable AI systems for Evaluating Arabic Pronunciations

  • 构建带标注的孤立阿拉伯字母语料库,用于语音识别研究
  • wav2vec 2.0 在该任务上仅达35%准确率,轻量模型提升至65%
  • 对抗训练使噪声下性能下降控制在9%以内,适合语言学习应用

现代阿拉伯语自动语音识别系统(如 wav2vec 2.0)在词级和句级转写上表现优异,但在孤立字母识别任务上表现不佳。该任务对语言学习、语音治疗和语音学研究至关重要,但因孤立字母缺乏协同发音线索、无词汇上下文且持续时间仅数百毫秒,导致识别困难。系统只能依赖可变声学特征,而阿拉伯语的强调辅音(喉化音)等发音在多数语言中无对应项,加剧了挑战。本研究构建了一个多样、带变音符号标注的孤立阿拉伯字母语料库,并验证 state-of-the-art wav2vec 2.0 模型在该数据集上准确率仅为35%。在 wav2vec 嵌入上训练轻量神经网络,准确率提升至65%。然而,添加小幅度扰动(epsilon = 0.05)后准确率降至32%。通过对抗训练,将噪声语音下的性能下降限制在9%以内,同时保持干净语音下的准确率。本文详述语料库、训练流程与评估协议,按需提供数据与代码以确保可复现性。最后提出未来工作方向:将此方法扩展至词级与句级框架,以持续支持精准发音评估。

原文摘要 · Abstract (English)

Modern Arabic ASR systems such as wav2vec 2.0 excel at word- and sentence-level transcription, yet struggle to classify isolated letters. In this study, we show that this phoneme-level task, crucial for language learning, speech therapy, and phonetic research, is challenging because isolated letters lack co-articulatory cues, provide no lexical context, and last only a few hundred milliseconds. Recogniser systems must therefore rely solely on variable acoustic cues, a difficulty heightened by Arabic's emphatic (pharyngealized) consonants and other sounds with no close analogues in many languages. This study introduces a diverse, diacritised corpus of isolated Arabic letters and demonstrates that state-of-the-art wav2vec 2.0 models achieve only 35% accuracy on it. Training a lightweight neural network on wav2vec embeddings raises performance to 65%. However, adding a small amplitude perturbation (epsilon = 0.05) cuts accuracy to 32%. To restore robustness, we apply adversarial training, limiting the noisy-speech drop to 9% while preserving clean-speech accuracy. We detail the corpus, training pipeline, and evaluation protocol, and release, on demand, data and code for reproducibility. Finally, we outline future work extending these methods to word- and sentence-level frameworks, where precise letter pronunciation remains critical.

语音识别阿拉伯语对抗训练发音评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。