arXiv:2505.02615cs.CLcs.SD2025-05被引 2

用深度学习自动评估英语二语水平,提升评分一致性。

Automatic Proficiency Assessment in L2 English Learners

  • 融合语音与文本分析,用CNN、ResNet和wav2vec 2.0模型综合评估。
  • 在EFCamDat、ANGLISH等数据集上达到接近人工评分的准确率。
  • 适合教育科技公司、语言测评平台快速部署自动化评分系统。

英语作为第二语言(L2)的水平通常由教师或专家主观评定,存在评分者间与评分者内差异。本文探索深度学习技术对L2英语能力进行综合评估,同时处理语音信号及其对应转录文本。我们比较了多种模型架构在语音表征分类中的表现,包括2D CNN、基于频率的CNN、ResNet以及预训练的wav2vec 2.0模型;同时在资源受限条件下,通过微调BERT模型实现文本层面的评分。此外,针对自发对话这一复杂任务,采用分离应用wav2vec 2.0和BERT模型的方法,处理长时音频与多说话人交互。在EFCamDat、ANGLISH数据集及一个私有数据集上的实验结果表明,尤其是预训练的wav2vec 2.0模型,在实现鲁棒的自动化L2英语能力评估方面展现出巨大潜力。

原文摘要 · Abstract (English)

Second language proficiency (L2) in English is usually perceptually evaluated by English teachers or expert evaluators, with the inherent intra- and inter-rater variability. This paper explores deep learning techniques for comprehensive L2 proficiency assessment, addressing both the speech signal and its correspondent transcription. We analyze spoken proficiency classification prediction using diverse architectures, including 2D CNN, frequency-based CNN, ResNet, and a pretrained wav2vec 2.0 model. Additionally, we examine text-based proficiency assessment by fine-tuning a BERT language model within resource constraints. Finally, we tackle the complex task of spontaneous dialogue assessment, managing long-form audio and speaker interactions through separate applications of wav2vec 2.0 and BERT models. Results from experiments on EFCamDat and ANGLISH datasets and a private dataset highlight the potential of deep learning, especially the pretrained wav2vec 2.0 model, for robust automated L2 proficiency evaluation.

语言评估深度学习语音分析自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。