arXiv:2505.04645cs.CLcs.LG2025-05被引 4

用ChatGPT评机械通气考试简答题,结果与人工评分差异大,不靠谱。

ChatGPT for automated grading of short answer questions in mechanical ventilation

  • 用标准评分框架让ChatGPT4o批改215名学生557份简答
  • 平均分低1.34分,个体评分一致性极差(ICC=0.086)
  • 尤其在分析题上错得离谱,不适合高风险考试评分

标准化的简答题测试广泛用于研究生教育。大型语言模型(LLMs)能模拟对话语言并理解自由文本回答,符合简答题评分量表要求,因此适用于自动化评分。本研究评估了ChatGPT 4o在机械通气在线课程(2020–2024)中对215名学生(共557份简答)的评分表现。将三组基于病例的题目及其标准评分提示和量表输入给ChatGPT,使用混合效应模型、方差成分分析、组内相关系数(ICC)、Cohen's kappa、Kendall's W及Bland-Altman统计法分析输出结果。ChatGPT评分系统性偏低,均值偏差(偏倚)为-1.34分(满分10分)。个体层面的一致性极差(ICC1 = 0.086),Cohen's kappa为-0.0786,表明无实质性一致性。方差成分分析显示五次ChatGPT会话间差异极小(G值=0.87),具内部一致性但严重偏离人工评分。评价性和分析性题目一致性最差,而清单型与指令型题项争议较小。超过60%的ChatGPT评分与人工评分差异超出高风险评估可接受范围,故建议谨慎使用LLM进行研究生课程评分。

原文摘要 · Abstract (English)

Standardised tests using short answer questions (SAQs) are common in postgraduate education. Large language models (LLMs) simulate conversational language and interpret unstructured free-text responses in ways aligning with applying SAQ grading rubrics, making them attractive for automated grading. We evaluated ChatGPT 4o to grade SAQs in a postgraduate medical setting using data from 215 students (557 short-answer responses) enrolled in an online course on mechanical ventilation (2020--2024). Deidentified responses to three case-based scenarios were presented to ChatGPT with a standardised grading prompt and rubric. Outputs were analysed using mixed-effects modelling, variance component analysis, intraclass correlation coefficients (ICCs), Cohen's kappa, Kendall's W, and Bland--Altman statistics. ChatGPT awarded systematically lower marks than human graders with a mean difference (bias) of -1.34 on a 10-point scale. ICC values indicated poor individual-level agreement (ICC1 = 0.086), and Cohen's kappa (-0.0786) suggested no meaningful agreement. Variance component analysis showed minimal variability among the five ChatGPT sessions (G-value = 0.87), indicating internal consistency but divergence from the human grader. The poorest agreement was observed for evaluative and analytic items, whereas checklist and prescriptive rubric items had less disagreement. We caution against the use of LLMs in grading postgraduate coursework. Over 60% of ChatGPT-assigned grades differed from human grades by more than acceptable boundaries for high-stakes assessments.

自动评分LLM医学教育ChatGPT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。