arXiv:2509.15926cs.CLcs.LG2025-09EMNLP被引 3

用置信度校准让作文评分模型更可靠,支持教师协作批改。

Beyond the Score: Uncertainty-Calibrated LLMs for Automated Essay Assessment

  • 用置信区间方法为评分模型加不确定度输出,保证90%覆盖
  • 在三个数据集上验证,评分集大小紧凑且准确率高
  • 适合需要可解释评分的教育场景,尤其高分考试应用

自动作文评分系统在部分公开基准上已接近人类评分一致度,但实际高风险考试中的应用仍受限。主要瓶颈是现有模型仅输出单一分数,缺乏置信度或解释。本文采用无需分布假设的共形预测方法,为任意分类器添加集合输出和严格的覆盖率保障。使用Llama-3 8B与Qwen-2.5 3B两个开源大模型,在ASAP、TOEFL11、Cambridge-FCE三个语料库上微调,并在90%风险水平下进行校准。通过一种新的不确定性感知准确率(UAcc)评估可靠性,该指标奖励模型在正确的同时保持输出简洁。据我们所知,这是首个将共形预测与UAcc结合用于作文评分的工作。校准后的模型始终满足覆盖率目标,且预测集规模紧凑,表明开源中等规模大模型已具备支持教师参与式评分的能力,未来可拓展至更大规模用户研究与系统扩展。

原文摘要 · Abstract (English)

Automated Essay Scoring (AES) systems now reach near human agreement on some public benchmarks, yet real-world adoption, especially in high-stakes examinations, remains limited. A principal obstacle is that most models output a single score without any accompanying measure of confidence or explanation. We address this gap with conformal prediction, a distribution-free wrapper that equips any classifier with set-valued outputs and formal coverage guarantees. Two open-source large language models (Llama-3 8B and Qwen-2.5 3B) are fine-tuned on three diverse corpora (ASAP, TOEFL11, Cambridge-FCE) and calibrated at a 90 percent risk level. Reliability is assessed with UAcc, an uncertainty-aware accuracy that rewards models for being both correct and concise. To our knowledge, this is the first work to combine conformal prediction and UAcc for essay scoring. The calibrated models consistently meet the coverage target while keeping prediction sets compact, indicating that open-source, mid-sized LLMs can already support teacher-in-the-loop AES; we discuss scaling and broader user studies as future work.

作文评分置信度校准共形预测LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。