arXiv:2505.21148cs.CLcs.SD2025-05被引 18

用大语言模型评估二语口语,效果超越传统方法。

Assessment of L2 Oral Proficiency using Speech Large Language Models

  • 用语音大模型直接评分,跳过中间环节损失信息
  • 在两个数据集上表现优于已有最佳模型
  • 跨说话人和任务评估时仍保持高准确率

随着二语英语学习者数量增加,自动口语评估系统需求上升。以往方法多依赖统计模型、文本编码器或自监督语音模型,但级联系统会丢失信息,端到端模型也有局限。随着多模态大语言模型的发展,本文探索其作为二语口语评分工具的潜力。通过对比回归与分类目标的不同训练策略,结果表明语音大模型全面超越先前基准,在两个数据集上表现更优。且经过预训练获得的音频理解能力使模型在跨说话人和跨任务评估中具备强泛化能力。

原文摘要 · Abstract (English)

The growing population of L2 English speakers has increased the demand for developing automatic graders for spoken language assessment (SLA). Historically, statistical models, text encoders, and self-supervised speech models have been utilised for this task. However, cascaded systems suffer from the loss of information, while E2E graders also have limitations. With the recent advancements of multi-modal large language models (LLMs), we aim to explore their potential as L2 oral proficiency graders and overcome these issues. In this work, we compare various training strategies using regression and classification targets. Our results show that speech LLMs outperform all previous competitive baselines, achieving superior performance on two datasets. Furthermore, the trained grader demonstrates strong generalisation capabilities in the cross-part or cross-task evaluation, facilitated by the audio understanding knowledge acquired during LLM pre-training.

语音大模型口语评估二语学习自动评分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。