arXiv:2603.11957cs.CL2026-03被引 1

让AI自动判卷时能识别自己是否靠谱,不靠谱的才交给人类。

CHiL(L)Grader: Calibrated Human-in-the-Loop Short-Answer Grading

  • 用校准置信度+选择性预测,只让高信心判断自动通过
  • 在三个数据集上自动评分35%-65%,准确率接近专家(QWK≥0.8)
  • 能随教学变化持续学习,适合需要安全可靠的评分场景

用大规模语言模型进行教育评估不仅需要准确性,还需识别预测是否可信。指令微调模型常过度自信,且随着课程演变,可靠性下降,导致在高风险场景中完全自主部署不安全。我们提出CHiL(L)Grader,首个将校准置信度融入人机协作流程的自动评分框架。通过后处理温度缩放、基于置信度的选择性预测和持续学习,该框架仅自动化高置信度预测,将不确定情况转给人工评分员,并适应不断变化的评分标准与未见题目。在三个短答案评分数据集上,CHiL(L)Grader可自动评分35%-65%的作答,达到专家水平(QWK ≥ 0.80)。接受与拒绝预测间的QWK差距达0.347,验证了置信度路由的有效性。每次纠错循环都使模型从教师反馈中提升评分能力。结果表明,不确定性量化是实现可靠AI辅助评分的关键。

原文摘要 · Abstract (English)

Scaling educational assessment with large language models requires not just accuracy, but the ability to recognize when predictions are trustworthy. Instruction-tuned models tend to be overconfident, and their reliability deteriorates as curricula evolve, making fully autonomous deployment unsafe in high-stakes settings. We introduce CHiL(L)Grader, the first automated grading framework that incorporates calibrated confidence estimation into a human-in-the-loop workflow. Using post-hoc temperature scaling, confidence-based selective prediction, and continual learning, CHiL(L)Grader automates only high-confidence predictions while routing uncertain cases to human graders, and adapts to evolving rubrics and unseen questions. Across three short-answer grading datasets, CHiL(L)Grader automatically scores 35-65% of responses at expert-level quality (QWK >= 0.80). A QWK gap of 0.347 between accepted and rejected predictions confirms the effectiveness of the confidence-based routing. Each correction cycle strengthens the model's grading capability as it learns from teacher feedback. These results show that uncertainty quantification is key for reliable AI-assisted grading.

自动评分置信度校准人机协作持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。