自动评分模型超越人类平均水平,且可解释。
Interpretable, Fairly Evaluated Automated L2 Speaking Assessment that Beats the Single-Human Ceiling and Why Pause Encoding Does Not Change LLM Fluency Scores

- 融合语音时序特征与大模型流畅性判断,实现可解释评分。
- 评分相关性达0.818,超过81%人类评分员,接近最优水平。
- 发现暂停标注不影响大模型流畅性评分,信号来自实际语速特征。
二语英语学习者极少有机会进行对话练习,口语也是最易引发焦虑的技能。这推动了自动化口语练习与评分市场的快速发展。但自动化评分只有在准确、可解释、公平且对标正确人类基准的前提下才可信。本文构建了一个基于特征+大模型的可解释混合模型,用于自发性二语对话评估。模型未对人类标签进行拟合,而是基于ICNALE全球评分档案(140篇演讲,约80名训练评分员,10项分析标准)进行评估。对130篇可用音频的二语演讲进行评分:确定性的De-Jong语音时序综合得分达到rho=0.764;与单一文本大模型流畅性判断融合后,与共识黄金标准的Spearman rho达0.818。该表现优于81%的人类评分员(中位数rho=0.73),接近最优水平,且达到可靠性校正后最大值(kappa_max=0.99)的约83%。融合相比纯复合模型提升+0.054(配对自助法95%置信区间[0.017, 0.108],不包含0),大模型提供粗粒度流畅性排序,由连续复合模型细化。此外,在控制条件下验证暂停编码无效,其影响被限制在±0.1 rho以内。保持大模型与学习者语词不变,仅改变暂停在提示中的写法时,行内标注位置不如整体暂停统计(-0.069,置信区间[-0.15, +0.08]),基于句中语义的暂停标准也无显著提升。流畅性信号来自实际语音时序特征,而非暂停的表达方式。所有结论均通过双学习者隔离方法、配对自助法置信区间、独白负控实验、经典测量的逐特征复现及按母语背景的公平性审计支持。
原文摘要 · Abstract (English)
Second-language (L2) English learners can rarely rehearse speaking with a partner. Speaking is also the most anxiety-laden skill. These gaps drive a fast-growing market for automated speaking practice and scoring. But an automated score is trustworthy only if it is accurate, interpretable, fair, and benchmarked against the right human bar. We build an interpretable feature-plus-LLM hybrid for spontaneous L2 dialogue. We evaluate it without ever fitting to the human labels, against the ICNALE Global Rating Archive: 140 speeches rated by ~80 trained raters on 10 analytic criteria. We score the 130 L2 speeches with usable audio. A deterministic De-Jong speech-timing composite reaches rho=0.764. Blended with a single text-LLM fluency judgment, it reaches Spearman rho=0.818 against the consensus gold. This agrees with the consensus better than 81% of the 80 individual trained raters: above the median rater (rho=0.73) and near the best, and at ~83% of the reliability-corrected maximum (kappa_max=0.99). The blend improves on the composite alone by +0.054 (paired-bootstrap 95% CI [0.017, 0.108], excludes 0); the LLM adds a coarse fluency ranking that the continuous composite refines. We also report a controlled null on pause encoding, bounded to effects below about +/-0.1 rho at this sample size. Holding the LLM and learner words fixed and varying only how pauses are written into the prompt, inline pause locations do not beat aggregate pause statistics (-0.069, CI [-0.15, +0.08]), and a grounded mid-clause criterion gives no reliable gain. The fluency signal comes from the measured speech-timing features, not from how pauses are written for the LLM. We back every claim with two agreeing learner-isolation methods, paired-bootstrap CIs, a monologue negative control, per-feature reproduction of classical measurements, and a per-L1 fairness audit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。