arXiv:2508.04105cs.AI2025-08

用GPT-4解释的多样性衡量评分不确定性,提升AI评分透明度。

Towards Transparent AI Grading: Semantic Entropy as a Signal for Human-AI Disagreement

  • 通过多条GPT-4解释的语义熵衡量评分分歧程度。
  • 在ASAP-SAS数据集上,熵值与人工评分者分歧相关性达0.68。
  • 适用于跨学科任务,尤其对需要解释的任务更敏感。

自动化评分系统可高效评估简答题,但常无法揭示评分结果的不确定性或潜在争议。本文提出以同一学生回答的多条GPT-4解释之间的语义熵作为人类评分者分歧的代理指标。通过基于蕴含相似性的归类并计算聚类间的熵,我们量化了理由的多样性,而不依赖最终得分。在ASAP-SAS数据集上的实验表明,语义熵与评分者分歧显著相关(r=0.68),在不同学科间有明显差异,并在需要解释性推理的任务中升高。研究结果表明,语义熵是一种可解释的不确定性信号,有助于构建更透明可信的AI辅助评分流程。

原文摘要 · Abstract (English)

Automated grading systems can efficiently score short-answer responses, yet they often fail to indicate when a grading decision is uncertain or potentially contentious. We introduce semantic entropy, a measure of variability across multiple GPT-4-generated explanations for the same student response, as a proxy for human grader disagreement. By clustering rationales via entailment-based similarity and computing entropy over these clusters, we quantify the diversity of justifications without relying on final output scores. We address three research questions: (1) Does semantic entropy align with human grader disagreement? (2) Does it generalize across academic subjects? (3) Is it sensitive to structural task features such as source dependency? Experiments on the ASAP-SAS dataset show that semantic entropy correlates with rater disagreement, varies meaningfully across subjects, and increases in tasks requiring interpretive reasoning. Our findings position semantic entropy as an interpretable uncertainty signal that supports more transparent and trustworthy AI-assisted grading workflows.

AI评分不确定性可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。