arXiv:2510.05162cs.CYcs.AI2025-10被引 3

用AI辅助批改微积分手写题,通过可信度过滤实现高精度与人工协作。

Artificial-Intelligence Grading Assistance for Handwritten Components of a Calculus Exam

  • 引入双层校验:分数阈值+项目反应理论风险评估
  • 严格过滤后AI准确率达人工水平,但70%题目仍需人工处理
  • 适合批量处理常规题,保留复杂题由教师判断

我们研究现代多模态大模型是否能在不损害评分有效性的前提下,大规模辅助批改开放题型的微积分手写作业。在一次大规模一年级考试中,学生的手写解答由GPT-5按同一名助教使用的评分细则进行评分,允许部分得分;助教评分作为真实标准。我们设计了一种人机协同过滤机制,结合部分分阈值与基于项目反应理论(2PL)的风险度量,衡量每个学生-题目对中AI评分与模型预期评分的偏差。未经过滤时,AI与助教评分一致性为中等,仅适用于低风险反馈,不适合高风险场景。通过置信度过滤,明确展现了工作量与质量的权衡:在更严格的设置下,AI达到人类准确率,但仍有约70%题目需人工完成。心理测量学特征受开放题部分低分值、评分节点少及答题区域与书写位置错位等因素制约。实际调整如略微提高权重、预留保护时间、增加可见子步骤、强化空间定位,可提升上限表现。总体而言,经过校准的置信度与保守路由策略,使AI能够可靠处理大量常规题,同时将模糊或教学价值高的响应留给专家判断。

原文摘要 · Abstract (English)

We investigate whether contemporary multimodal LLMs can assist with grading open-ended calculus at scale without eroding validity. In a large first-year exam, students' handwritten work was graded by GPT-5 against the same rubric used by teaching assistants (TAs), with fractional credit permitted; TA rubric decisions served as ground truth. We calibrated a human-in-the-loop filter that combines a partial-credit threshold with an Item Response Theory (2PL) risk measure based on the deviation between the AI score and the model-expected score for each student-item. Unfiltered AI-TA agreement was moderate, adequate for low-stakes feedback but not for high-stakes use. Confidence filtering made the workload-quality trade-off explicit: under stricter settings, AI delivered human-level accuracy, but also left roughly 70% of the items to be graded by humans. Psychometric patterns were constrained by low stakes on the open-ended portion, a small set of rubric checkpoints, and occasional misalignment between designated answer regions and where work appeared. Practical adjustments such as slightly higher weight and protected time, a few rubric-visible substeps, stronger spatial anchoring should raise ceiling performance. Overall, calibrated confidence and conservative routing enable AI to reliably handle a sizable subset of routine cases while reserving expert judgment for ambiguous or pedagogically rich responses.

AI批改多模态模型教育评测自动化评分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。