用多目标学习实现整段口语评分,提升语言评估准确性。
Session-Level Spoken Language Assessment with a Multimodal Foundation Model via Multi-Target Learning
- 联合学习整体能力和特质水平的评分目标,避免人工特征。
- 基于冻结的Whisper模型进行声学校准,支持全会话分析。
- 在Speak & Improve数据集上超越现有系统,适合实际教学应用。
口语能力评估(SLA)通过自发口语估计学习者的口述能力。随着第二语言英语学习者数量增长,可靠SLA成为计算机辅助语言学习(CALL)的关键组成部分。现有方法多采用级联流水线,易产生误差传播;或使用短音频窗口的端到端模型,可能遗漏语篇级证据。本文提出一种新型多模态基础模型方法,通过单次处理实现会话级评估。该方法结合多目标学习与冻结的Whisper ASR模型作为语音先验,实现声学感知校准,无需手工特征即可联合学习整体与特质级的SLA目标。通过统一处理整个回答会话,模型在预测整体口语水平方面表现优异。在Speak & Improve基准上的实验表明,该方法优于先前最先进级联系统,并展现出稳健的跨参与者泛化能力,生成一个紧凑可部署的评分器,专为CALL应用设计。
原文摘要 · Abstract (English)
Spoken Language Assessment (SLA) estimates a learner's oral proficiency from spontaneous speech. The growing population of L2 English speakers has intensified the demand for reliable SLA, a critical component of Computer Assisted Language Learning (CALL). Existing efforts often rely on cascaded pipelines, which are prone to error propagation, or end-to-end models that often operate on a short audio window, which might miss discourse-level evidence. This paper introduces a novel multimodal foundation model approach that performs session-level evaluation in a single pass. Our approach couples multi-target learning with a frozen, Whisper ASR model-based speech prior for acoustic-aware calibration, allowing for jointly learning holistic and trait-level objectives of SLA without resorting to handcrafted features. By coherently processing the entire response session of an L2 speaker, the model excels at predicting holistic oral proficiency. Experiments conducted on the Speak & Improve benchmark demonstrate that our proposed approach outperforms the previous state-of-the-art cascaded system and exhibits robust cross-part generalization, producing a compact deployable grader that is tailored for CALL applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。