arXiv:2510.27131cs.LG2025-10被引 1

用大模型生成的推理过程提升作文自动评分,尤其改善了低分段准确率。

Exploring the Utilities of the Rationales from Large Language Models to Enhance Automated Essay Scoring

  • 利用GPT-4.1和GPT-5生成作文推理理由,与原文评分对比。
  • 低分(0分)的F1分数显著提升,缓解类别不平衡问题。
  • 融合原文与推理评分模型,达到0.870的QWK,优于文献结果。

本研究探索了GPT-4.1和GPT-5生成的推理过程在自动化作文评分中的效用,基于2012年Kaggle ASAP数据集中的Prompt 6作文。对比了基于原文的评分与基于推理理由的评分。总体上,原文评分表现更优,具有更高的加权二次κ系数(QWK)。然而,基于推理的评分在0分这一代表性不足的类别上,其F1分数更高,有效缓解了类别不平衡问题。将多个原文评分模型集成后,在各分数层级及整体上均提升了准确率。原文与单个推理评分模型的集成效果相当。进一步融合原文与两个推理评分模型,实现了最佳性能,QWK达0.870,优于文献中报告的0.848。

原文摘要 · Abstract (English)

This study explored the utilities of rationales generated by GPT-4.1 and GPT-5 in automated scoring using Prompt 6 essays from the 2012 Kaggle ASAP data. Essay-based scoring was compared with rationale-based scoring. The study found in general essay-based scoring performed better than rationale-based scoring with higher Quadratic Weighted Kappa (QWK). However, rationale-based scoring led to higher scoring accuracy in terms of F1 scores for score 0 which had less representation due to class imbalance issues. The ensemble modeling of essay-based scoring models increased the scoring accuracy at both specific score levels and across all score levels. The ensemble modeling of essay-based scoring and each of the rationale-based scoring performed about the same. Further ensemble of essay-based scoring and both rationale-based scoring yielded the best scoring accuracy with QWK of 0.870 compared with 0.848 reported in literature.

自动评分大模型推理作文评估集成学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。