用大模型提升数学答题评分精度,解决难题案例的误判问题。
Learning to Love Edge Cases in Formative Math Assessment: Using the AMMORE Dataset and Chain-of-Thought Prompting to Improve Grading Accuracy
- 通过思维链提示让大模型精准识别复杂答题
- 难题评分准确率从98.7%提升至99.9%
- 改善学生掌握程度评估,误判率下降至2.6%
本文引入AMMORE数据集,包含来自非洲多国学习平台Rori的53,000个数学开放作答题对,并开展两项实验,评估大语言模型(LLM)在评分高难度学生作答中的表现。实验一采用零样本、少样本及思维链提示等方法,针对规则模型无法准确评分的1%边缘案例进行评分,发现思维链提示最佳,可将这些案例的准确率提升至92%,使整体评分准确率从98.7%提升至99.9%。实验二通过将最优LLM评分结果输入贝叶斯知识追踪(BKT)模型,分析其对学生掌握状态估计的影响,发现微小的单题评分改进可显著改变掌握度判断:规则模型误判6.9%学生的掌握状态,而使用LLM思维链方法后,该比例降至2.6%。结果表明,LLM在中小学数学形成性评价中具有重要应用潜力,有助于推动开放题更广泛使用。
原文摘要 · Abstract (English)
This paper introduces AMMORE, a new dataset of 53,000 math open-response question-answer pairs from Rori, a learning platform used by students in several African countries and conducts two experiments to evaluate the use of large language models (LLM) for grading particularly challenging student answers. The AMMORE dataset enables various potential analyses and provides an important resource for researching student math acquisition in understudied, real-world, educational contexts. In experiment 1 we use a variety of LLM-driven approaches, including zero-shot, few-shot, and chain-of-thought prompting, to grade the 1% of student answers that a rule-based classifier fails to grade accurately. We find that the best-performing approach -- chain-of-thought prompting -- accurately scored 92% of these edge cases, effectively boosting the overall accuracy of the grading from 98.7% to 99.9%. In experiment 2, we aim to better understand the consequential validity of the improved grading accuracy, by passing grades generated by the best-performing LLM-based approach to a Bayesian Knowledge Tracing (BKT) model, which estimated student mastery of specific lessons. We find that relatively modest improvements in model accuracy at the individual question level can lead to significant changes in the estimation of student mastery. Where the rules-based classifier currently used to grade student, answers misclassified the mastery status of 6.9% of students across their completed lessons, using the LLM chain-of-thought approach this misclassification rate was reduced to 2.6% of students. Taken together, these findings suggest that LLMs could be a valuable tool for grading open-response questions in K-12 mathematics education, potentially enabling encouraging wider adoption of open-ended questions in formative assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。