对比10个AI与人工评分,发现3个AI更准更稳。
Comparing Human and AI Rater Effects Using the Many-Facet Rasch Model
- 用多面拉斯模型量化评分偏差,比较AI与人类评分者。
- ChatGPT 4o等3款AI评分准确率高,一致性好。
- 适合教育评估自动化研究者参考,尤其关注评分公平性。
大型语言模型(LLMs)在低门槛测评中被广泛用于自动评分,以支持学习与教学。但在实际应用前,需收集实证证据,明确哪个模型评分最可靠、产生的评分者效应最小。本研究对比了十种LLMs(ChatGPT 3.5、ChatGPT 4、ChatGPT 4o、OpenAI o1、Claude 3.5 Sonnet、Gemini 1.5、Gemini 1.5 Pro、Gemini 2.0、DeepSeek V3 和 DeepSeek R1)与人类专家评分者,在两种写作任务上的评分表现。通过加权二次卡帕系数(Quadratic Weighted Kappa)评估整体与分析评分的准确性;通过克隆巴赫α系数(Cronbach Alpha)评估同一评分者在不同提示下的内部一致性;并采用多面拉斯模型(Many-Facet Rasch Model)评估和比较各评分者的评分效应。结果总体表明,ChatGPT 4o、Gemini 1.5 Pro 和 Claude 3.5 Sonnet 在评分准确性、评分者可靠性及评分效应方面均表现优异。
原文摘要 · Abstract (English)
Large language models (LLMs) have been widely explored for automated scoring in low-stakes assessment to facilitate learning and instruction. Empirical evidence related to which LLM produces the most reliable scores and induces least rater effects needs to be collected before the use of LLMs for automated scoring in practice. This study compared ten LLMs (ChatGPT 3.5, ChatGPT 4, ChatGPT 4o, OpenAI o1, Claude 3.5 Sonnet, Gemini 1.5, Gemini 1.5 Pro, Gemini 2.0, as well as DeepSeek V3, and DeepSeek R1) with human expert raters in scoring two types of writing tasks. The accuracy of the holistic and analytic scores from LLMs compared with human raters was evaluated in terms of Quadratic Weighted Kappa. Intra-rater consistency across prompts was compared in terms of Cronbach Alpha. Rater effects of LLMs were evaluated and compared with human raters using the Many-Facet Rasch model. The results in general supported the use of ChatGPT 4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet with high scoring accuracy, better rater reliability, and less rater effects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。