评测大模型自动评分中的不确定性,找出可靠度量方法。
How Uncertain Is the Grade? A Benchmark of Uncertainty Metrics for LLM-Based Automatic Assessment
- 对比多种不确定性度量方法在自动评分中的表现。
- 发现模型类型和生成策略显著影响不确定性估计。
- 为构建更可信的智能评分系统提供实证依据。
大型语言模型(LLMs)正重塑教育领域的自动评估格局。尽管其在适应多样题型和灵活输出方面具有优势,但其固有的概率特性也带来了输出不确定性问题。不准确或未校准的不确定性估计可能导致教学干预失稳,影响学生学习。为此,我们系统性地评测了多种不确定性量化方法在大模型自动评分中的表现。通过在多个评估数据集、模型家族和生成控制设置下的综合分析,我们刻画了大模型在评分任务中表现出的不确定性模式,并评估了不同度量方法的优劣。研究揭示了模型家族、评估任务和解码策略对不确定性估计的影响,为未来开发更可靠、可信赖的不确定性感知评分系统提供了关键洞见。
原文摘要 · Abstract (English)
The rapid rise of large language models (LLMs) is reshaping the landscape of automatic assessment in education. While these systems demonstrate substantial advantages in adaptability to diverse question types and flexibility in output formats, they also introduce new challenges related to output uncertainty, stemming from the inherently probabilistic nature of LLMs. Output uncertainty is an inescapable challenge in automatic assessment, as assessment results often play a critical role in informing subsequent pedagogical actions, such as providing feedback to students or guiding instructional decisions. Unreliable or poorly calibrated uncertainty estimates can lead to unstable downstream interventions, potentially disrupting students' learning processes and resulting in unintended negative consequences. To systematically understand this challenge and inform future research, we benchmark a broad range of uncertainty quantification methods in the context of LLM-based automatic assessment. Although the effectiveness of these methods has been demonstrated in many tasks across other domains, their applicability and reliability in educational settings, particularly for automatic grading, remain underexplored. Through comprehensive analyses of uncertainty behaviors across multiple assessment datasets, LLM families, and generation control settings, we characterize the uncertainty patterns exhibited by LLMs in grading scenarios. Based on these findings, we evaluate the strengths and limitations of different uncertainty metrics and analyze the influence of key factors, including model families, assessment tasks, and decoding strategies, on uncertainty estimates. Our study provides actionable insights into the characteristics of uncertainty in LLM-based automatic assessment and lays the groundwork for developing more reliable and effective uncertainty-aware grading systems in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。