用AI辅助批改微积分手写题,通过可信度过滤实现高精度与人工协作。
Artificial-Intelligence Grading Assistance for Handwritten Components of a Calculus Exam
- 引入双层校验:分数阈值+项目反应理论风险评估
- 严格过滤后AI准确率达人工水平,但70%题目仍需人工处理
- 适合批量处理常规题,保留复杂题由教师判断
我们研究现代多模态大模型是否能在不损害评分有效性的前提下,大规模辅助批改开放题型的微积分手写作业。在一次大规模一年级考试中,学生的手写解答由GPT-5按同一名助教使用的评分细则进行评分,允许部分得分;助教评分作为真实标准。我们设计了一种人机协同过滤机制,结合部分分阈值与基于项目反应理论(2PL)的风险度量,衡量每个学生-题目对中AI评分与模型预期评分的偏差。未经过滤时,AI与助教评分一致性为中等,仅适用于低风险反馈,不适合高风险场景。通过置信度过滤,明确展现了工作量与质量的权衡:在更严格的设置下,AI达到人类准确率,但仍有约70%题目需人工完成。心理测量学特征受开放题部分低分值、评分节点少及答题区域与书写位置错位等因素制约。实际调整如略微提高权重、预留保护时间、增加可见子步骤、强化空间定位,可提升上限表现。总体而言,经过校准的置信度与保守路由策略,使AI能够可靠处理大量常规题,同时将模糊或教学价值高的响应留给专家判断。
原文摘要 · Abstract (English)
We investigate whether contemporary multimodal LLMs can assist with grading open-ended calculus at scale without eroding validity. In a large first-year exam, students' handwritten work was graded by GPT-5 against the same rubric used by teaching assistants (TAs), with fractional credit permitted; TA rubric decisions served as ground truth. We calibrated a human-in-the-loop filter that combines a partial-credit threshold with an Item Response Theory (2PL) risk measure based on the deviation between the AI score and the model-expected score for each student-item. Unfiltered AI-TA agreement was moderate, adequate for low-stakes feedback but not for high-stakes use. Confidence filtering made the workload-quality trade-off explicit: under stricter settings, AI delivered human-level accuracy, but also left roughly 70% of the items to be graded by humans. Psychometric patterns were constrained by low stakes on the open-ended portion, a small set of rubric checkpoints, and occasional misalignment between designated answer regions and where work appeared. Practical adjustments such as slightly higher weight and protected time, a few rubric-visible substeps, stronger spatial anchoring should raise ceiling performance. Overall, calibrated confidence and conservative routing enable AI to reliably handle a sizable subset of routine cases while reserving expert judgment for ambiguous or pedagogically rich responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。