用自然语言描述评估二语口语,让大模型像真人一样打分。
Natural Language-based Assessment of L2 Oral Proficiency using LLMs
- 用Qwen 2.5 72B模型零样本理解语言描述进行评分
- 在S&I语料上表现优于专用BERT模型,接近语音大模型
- 评分逻辑透明可解释,适用于多语言和不同任务
自然语言评估(NLA)是一种利用原始为人类考官设计的‘能做’描述语来评估第二语言能力的方法,旨在检验大语言模型(LLMs)是否能以与人类评估相当的方式理解并应用这些描述。本文探索使用开源大模型Qwen 2.5 72B,在零样本设置下对公开的S&I语料库中的口语回答进行评估。结果显示,仅依赖文本信息的方法性能具有竞争力:虽未超越微调过的语音大模型,但优于专门为此任务训练的BERT模型。NLA在任务不匹配场景中尤为有效,具备跨数据类型和语言的泛化能力,且因基于清晰可解释的通用语言描述而具有更高可解释性。
原文摘要 · Abstract (English)
Natural language-based assessment (NLA) is an approach to second language assessment that uses instructions - expressed in the form of can-do descriptors - originally intended for human examiners, aiming to determine whether large language models (LLMs) can interpret and apply them in ways comparable to human assessment. In this work, we explore the use of such descriptors with an open-source LLM, Qwen 2.5 72B, to assess responses from the publicly available S&I Corpus in a zero-shot setting. Our results show that this approach - relying solely on textual information - achieves competitive performance: while it does not outperform state-of-the-art speech LLMs fine-tuned for the task, it surpasses a BERT-based model trained specifically for this purpose. NLA proves particularly effective in mismatched task settings, is generalisable to other data types and languages, and offers greater interpretability, as it is grounded in clearly explainable, widely applicable language descriptors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。