用大模型辅助数学考试评分,提升批改效率与公平性。
LLMs as Teaching Assistants for Mathematics Exam Grading: Reliability, and Practical Usability
- 采用宽松部分得分提示,降低大模型评分误差
- ChatGPT 5.5 Thinking(宽松)题级误差最低(MAE 1.87)
- Gemini 3.1 Pro Extended(基准)总分相关性最强(0.58)
开放式数学考试能有效评估推理、证明构建、算法思维和中间步骤表达,但大规模批改困难,需教师一致应用部分分评分标准并提供助于纠正误解的反馈。本文评估六种主流大语言模型(Gemini 3.1 Pro Extended、Gemini 3.5 Flash、ChatGPT 5.5 Pro Extended、ChatGPT 5.5 Thinking、Claude Pro Opus 4.7、Claude Sonnet 4.6)在本科离散数学考试中的评分助手表现。比较两种评分策略:基准(强调明确证据与完整论证)和宽松(修正初始严苛扣分)。通过平均绝对误差(MAE)、均方根误差(RMSE)、归一化均方根误差、皮尔逊相关系数和精确一致性衡量与人工评分的一致性。结果显示,宽松部分分提示可降低所有模型的平均题级误差。ChatGPT 5.5 Thinking(宽松)题级误差最低(MAE 1.87,RMSE 2.53),Gemini 3.1 Pro Extended(宽松)总分误差最低(MAE 8.00,RMSE 10.66)。但总分皮尔逊相关系数最高(0.58)出现在Gemini 3.1 Pro Extended(基准)策略下,表明分数校准与排名保持是不同目标。同时报告了实际可用性观察。
原文摘要 · Abstract (English)
Open-ended mathematics exams are valuable because they assess reasoning, proof construction, algorithmic thinking, and communication of intermediate steps. They are also difficult to grade at scale because instructors must apply partial-credit rubrics consistently while giving feedback that helps students repair misconceptions. This paper evaluates six contemporary large language model (LLM) configurations, Gemini 3.1 Pro Extended, Gemini 3.5 Flash, ChatGPT 5.5 Pro Extended, ChatGPT 5.5 Thinking, Claude Pro Opus 4.7, and Claude Sonnet 4.6, as grading assistants for an undergraduate discrete mathematics examination. The study compares two grading policies. The BASELINE policy uses a stricter rubric-following prompt that emphasizes explicit evidence and complete justification. The LIBERAL policy was added after preliminary grading showed that the baseline condition sometimes applied harsh point deductions and failed to recognize valid partial reasoning. Agreement with human grading is measured at both the question and exam-total levels using mean absolute error, root mean squared error, normalized root mean squared error, Pearson correlation, and exact agreement. The results show that liberal partial-credit prompting reduces average question-level error for every evaluated model family. ChatGPT 5.5 Thinking (LIBERAL) has the lowest average question-level MAE (1.87) and RMSE (2.53), while Gemini 3.1 Pro Extended (LIBERAL) has the lowest total-score MAE (8.00) and RMSE (10.66). However, the strongest total-score Pearson correlation occurs under Gemini 3.1 Pro Extended (BASELINE) at 0.58, showing that point calibration and rank preservation remain distinct goals. We also report practical usability observations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。