arXiv:2409.14494cs.CLcs.LG2024-09被引 13

用持续预训练提升语音识别在教室环境的抗噪能力

CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments

  • 对Wav2vec2.0进行教室场景持续预训练
  • 词错误率降低超10%,显著提升抗噪性能
  • 适合教育AI、智能助教系统开发者参考

构建在教室环境下稳健可靠的自动语音识别(ASR)系统,对发展辅助教师与学生的AI工具至关重要。本文研究了持续预训练(CPT)在将Wav2vec2.0适配至教室领域时的有效性。结果表明,CPT是一种强大手段,可使基于Wav2vec2.0的模型词错误率(WER)降低超过10%。具体而言,该方法提升了模型对不同噪声、麦克风及教室环境的鲁棒性。

原文摘要 · Abstract (English)

Creating Automatic Speech Recognition (ASR) systems that are robust and resilient to classroom conditions is paramount to the development of AI tools to aid teachers and students. In this work, we study the efficacy of continued pretraining (CPT) in adapting Wav2vec2.0 to the classroom domain. We show that CPT is a powerful tool in that regard and reduces the Word Error Rate (WER) of Wav2vec2.0-based models by upwards of 10%. More specifically, CPT improves the model's robustness to different noises, microphones and classroom conditions.

语音识别教室场景持续预训练降噪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。