arXiv:2509.23504cs.CL2025-09被引 1

基于双阶段训练的阿拉伯语语音转音素系统,获2025年Iqra'Eval竞赛第一名。

AraS2P: Arabic Speech-to-Phonemes System

  • 用阿拉伯语语音-音素数据进行任务自适应预训练,再微调优化。
  • 通过合成诵读数据增强,音素级误读检测性能显著提升。
  • 适合研究低资源语言音素识别与语音错误检测的学者。

本文介绍我们提交至Iqra'Eval 2025共享任务的AraS2P语音转音素系统。我们采用两阶段训练策略对Wav2Vec2-BERT进行适配:第一阶段在大规模阿拉伯语语音-音素数据集上进行任务自适应持续预训练,这些数据通过MSA Phonetiser将阿拉伯语文本转换生成;第二阶段在官方共享任务数据上微调模型,并引入来自XTTS-v2合成诵读的额外数据增强,包含不同阿亚段落、说话人嵌入及文本扰动,以模拟可能的人类错误。该系统在官方排行榜上排名第一,表明音素感知预训练结合针对性数据增强,在音素级误读检测任务中表现优异。

原文摘要 · Abstract (English)

This paper describes AraS2P, our speech-to-phonemes system submitted to the Iqra'Eval 2025 Shared Task. We adapted Wav2Vec2-BERT via Two-Stage training strategy. In the first stage, task-adaptive continue pretraining was performed on large-scale Arabic speech-phonemes datasets, which were generated by converting the Arabic text using the MSA Phonetiser. In the second stage, the model was fine-tuned on the official shared task data, with additional augmentation from XTTS-v2-synthesized recitations featuring varied Ayat segments, speaker embeddings, and textual perturbations to simulate possible human errors. The system ranked first on the official leaderboard, demonstrating that phoneme-aware pretraining combined with targeted augmentation yields strong performance in phoneme-level mispronunciation detection.

语音识别音素识别阿拉伯语数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。