用大模型评估教师教学知识,效率高但也有评分偏差。
Using Large Language Models to Assess Teachers' Pedagogical Content Knowledge
- 用大模型、人工和机器学习三种方式评分对比偏差来源。
- 大模型比人工更宽松,对情境变化不敏感,但整体偏差可控。
- 适合教育评估自动化研究者与智能评分系统设计者参考。
通过表现性任务评估教师的学科教学知识(PCK)耗时费力。虽然大语言模型(LLMs)为自动评分提供了新可能,但尚不清楚其是否像传统机器学习(ML)或人工评分一样引入无关构念的变异(CIV)。本研究聚焦视频类问答任务中两个PCK子维度——分析学生思维、评价教师回应性,考察场景差异、评分者严苛度、对场景敏感度三类CIV来源。采用广义线性混合模型(GLMMs)比较人工评分员、监督式机器学习模型与大模型三者的方差成分与评分模式。结果表明,任务层面的场景差异影响较小,而评分者相关因素贡献了主要的CIV,尤其在更依赖解释的Task II中更为明显。监督式机器学习模型最为严苛且最不敏感,而大模型则最为宽松。研究显示,大模型虽提升评分效率,但仍引入类似人工评分的构念无关变异,程度与监督式机器学习不同。讨论了评分员培训、自动化评分设计及模型可解释性未来研究的意义。
原文摘要 · Abstract (English)
Assessing teachers' pedagogical content knowledge (PCK) through performance-based tasks is both time and effort-consuming. While large language models (LLMs) offer new opportunities for efficient automatic scoring, little is known about whether LLMs introduce construct-irrelevant variance (CIV) in ways similar to or different from traditional machine learning (ML) and human raters. This study examines three sources of CIV -- scenario variability, rater severity, and rater sensitivity to scenario -- in the context of video-based constructed-response tasks targeting two PCK sub-constructs: analyzing student thinking and evaluating teacher responsiveness. Using generalized linear mixed models (GLMMs), we compared variance components and rater-level scoring patterns across three scoring sources: human raters, supervised ML, and LLM. Results indicate that scenario-level variance was minimal across tasks, while rater-related factors contributed substantially to CIV, especially in the more interpretive Task II. The ML model was the most severe and least sensitive rater, whereas the LLM was the most lenient. These findings suggest that the LLM contributes to scoring efficiency while also introducing CIV as human raters do, yet with varying levels of contribution compared to supervised ML. Implications for rater training, automated scoring design, and future research on model interpretability are discussed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。