arXiv:2603.16889cs.CLcs.AI2026-03中稿 · LREC 2026被引 3

用评分标准引导模型评估二语口语,提升准确性和可信度。

Rubric-Guided Fine-tuning of SpeechLLMs for Multi-Aspect, Multi-Rater L2 Reading-Speech Assessment

  • 基于评分标准设计多维度评估框架,融合准确性、流利度和语调。
  • 通过不确定性校准,使评分结果与人工评判高度一致。
  • 适合需要可解释性评分的教育AI场景,如语言考试系统。

可靠的多方面、多评分者第二语言(L2)口语自动化评估仍是核心挑战,因大型语音语言模型(SpeechLLMs)常难以匹配人类评分者的细微差异。为此,我们提出一种基于评分标准的推理框架,显式编码准确性、流利度和语调等多维度人类评估标准,并校准模型不确定性以捕捉自然评分变异性。我们使用多评分者人工判断对Qwen2-Audio-7B-Instruct模型进行微调,采用基于共形校准的不确定性校准回归方法,生成可解释的置信区间。该方法在高斯不确定性建模与共形校准支持下,实现了与人工评分的最佳一致性,显著优于回归与分类基线。模型能稳定评估流利度与语调,但准确性评估仍具挑战。结果表明,基于评分标准与不确定性校准的推理路径,为构建可信且可解释的SpeechLLM语音评估提供了原则性方案。

原文摘要 · Abstract (English)

Reliable and interpretable automated assessment of second-language (L2) speech remains a central challenge, as large speech-language models (SpeechLLMs) often struggle to align with the nuanced variability of human raters. To address this, we introduce a rubric-guided reasoning framework that explicitly encodes multi-aspect human assessment criteria: accuracy, fluency, and prosody, while calibrating model uncertainty to capture natural rating variability. We fine-tune the Qwen2-Audio-7B-Instruct model using multi-rater human judgments and develop an uncertainty-calibrated regression approach supported by conformal calibration for interpretable confidence intervals. Our Gaussian uncertainty modeling and conformal calibration approach achieves the strongest alignment with human ratings, outperforming regression and classification baselines. The model reliably assesses fluency and prosody while highlighting the inherent difficulty of assessing accuracy. Together, these results demonstrate that rubric-guided, uncertainty-calibrated reasoning offers a principled path toward trustworthy and explainable SpeechLLM-based speech assessment.

语音评估评分标准不确定性校准二语学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。