arXiv:2605.00200cs.CL2026-05中稿 · the 27th Internati…被引 1

提升AI评分可信度,让系统自知何时该听人话。

Confidence Estimation in Automatic Short Answer Grading with LLMs

  • 融合模型自身判断与学生作答数据的不确定性
  • 混合方法使评分置信度更可靠,选择性评分效果更好
  • 适合教育评估中需人机协作的场景

使用生成式大语言模型(LLMs)进行自动短答案评分近年来展现出强大性能,无需特定任务微调即可实现,还能生成教学反馈。然而,基于LLM的评分仍不完美,可靠的置信度估计对教育决策中的人机协作至关重要。本文研究了在无任务微调条件下,通过结合模型自身置信信号与数据集导出的不确定性来实现置信度估计。系统比较了三种模型置信度估计策略:语义化、潜在空间和一致性方法,发现仅靠模型置信度无法有效捕捉评分中的不确定性。为此,提出一种混合置信度框架,将模型置信信号与显式的数据集级随机不确定性相结合。该不确定性通过嵌入学生答案并聚类,量化簇内异质性来实现。实验表明,所提混合置信度测量方法比单一来源方法更可靠,并提升了选择性评分性能。本工作推动了人在回路中的可信度感知型智能评分,支持更可信赖的AI辅助教育评估系统。

原文摘要 · Abstract (English)

Automatic Short Answer Grading (ASAG) with generative large language models (LLMs) has recently demonstrated strong performance without task-specific fine-tuning, while also enabling the generation of synthetic feedback for educational assessment. Despite these advances, LLM-based grading remains imperfect, making reliable confidence estimates essential for safe and effective human-AI collaboration in educational decision-making. In this work, we investigate confidence estimation for ASAG with LLMs by jointly considering model-based confidence signals and dataset-derived uncertainty. We systematically compare three model-based confidence estimation strategies, namely verbalizing, latent, and consistency-based confidence estimation, and show that model-based confidence alone is insufficient to reliably capture uncertainty in ASAG. To address this limitation, we propose a hybrid confidence framework that integrates model-based confidence signals with an explicit estimate of dataset-derived aleatoric uncertainty. Aleatoric uncertainty is operationalized by clustering semantically embedded student responses and quantifying within-cluster heterogeneity. Our results demonstrate that the proposed hybrid confidence measure yields more reliable confidence estimates and improves selective grading performance compared to single-source approaches. Overall, this work advances confidence-aware LLM-based grading for human-in-the-loop assessment, supporting more trustworthy AI-assisted educational assessment systems.

AI评分置信度估计教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。