用清晰的评分标准让AI更准评物理答题,尤其适合高分和低分题。
Designing Reliable LLM-Assisted Rubric Scoring for Constructed Responses: Evidence from Physics Exams
- 用分项技能清单式评分表,提升AI打分一致性。
- 高分与低分题的AI评分准确率接近人工,中等题易出错。
- 评分标准比提示格式更重要,温度影响小。
STEM考试中的学生作答常为手写,包含符号表达、计算和图表,格式差异大,评分耗时且易因评分者不一致而失准,尤其在需部分赋分时。本研究使用GPT-4o评估了本科物理问答题的AI辅助评分可靠性。20份真实手写答卷由四位教师在两轮中评分,并由AI基于不同粒度的技能型评分标准进行评分,系统调整提示格式与温度参数。总体上,人类与AI在总分上的一致性相当于人工评分者间一致性,且在高分与低分题上表现最佳,中等水平题(涉及部分或模糊推理)则下降。逐项分析显示,对明确概念技能的判断一致性更高,而对复杂过程判断较弱。更细粒度的检查表式评分表优于整体式评分。结果表明,可靠AI评分主要依赖清晰结构化的评分标准,提示格式作用次之,温度影响较小。研究为在STEM场景中通过技能型评分标准与可控模型设置实现可靠AI评分提供了可迁移的设计建议。
原文摘要 · Abstract (English)
Student responses in STEM assessments are often handwritten and combine symbolic expressions, calculations, and diagrams, creating substantial variation in format and interpretation. Despite their importance for evaluating students' reasoning, such responses are time-consuming to score and prone to rater inconsistency, particularly when partial credit is required. Recent advances in large language models (LLMs) have increased attention to AI-assisted scoring, yet evidence remains limited regarding how rubric design and LLM configurations influence reliability across performance levels. This study examined the reliability of AI-assisted scoring of undergraduate physics constructed responses using GPT-4o. Twenty authentic handwritten exam responses were scored across two rounds by four instructors and by the AI model using skill-based rubrics with differing levels of analytic granularity. Prompting format and temperature settings were systematically varied. Overall, human-AI agreement on total scores was comparable to human inter-rater reliability and was highest for high- and low-performing responses, but declined for mid-level responses involving partial or ambiguous reasoning. Criterion-level analyses showed stronger alignment for clearly defined conceptual skills than for extended procedural judgments. A more fine-grained, checklist-based rubric improved consistency relative to holistic scoring. These findings indicate that reliable AI-assisted scoring depends primarily on clear, well-structured rubrics, while prompting format plays a secondary role and temperature has relatively limited impact. More broadly, the study provides transferable design recommendations for implementing reliable LLM-assisted scoring in STEM contexts through skill-based rubrics and controlled LLM settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。