arXiv:2510.22968cs.CLcs.AI2025-10

用定制大模型评估教学品质,效果媲美甚至超过人类专家。

Measuring Teaching with LLMs

  • 基于句子嵌入构建专用LLM,更适配课堂长文本分析。
  • 模型评分与人类专家相关性超0.65,部分指标超越人类平均。
  • 能捕捉整体课程特征,适合教师发展反馈,但个体细节仍有不足。

客观且可扩展的教学质量测量是教育领域的长期挑战。尽管大语言模型(LLMs)具有潜力,通用模型在应用复杂、真实的课堂观察工具时表现不稳定。本文采用基于句子级嵌入的定制化LLM架构,该架构比传统子词分词更适合长篇、解释性课堂转录文本。在数据高效训练策略下,系统评估了五种不同句子嵌入方法,结果表明这些专用模型在专家人工评分上达到甚至超过人类水平,平均人-人评分相关性高于0.65。通过分析标注上下文窗口发现,更接近人类判断的先进模型将更多分数差异归因于课程层面特征而非孤立语句,质疑了单轮标注范式的充分性。此外,聚合模型评分与教师增值效应一致,表明其捕捉到了对学生学习有影响的特征;但个体题目层面未呈现此趋势,说明模型虽学到有用信号,尚未实现完全泛化。本研究建立了一种可行且强大的人工智能驱动教学评估新方法,为教师发展提供可扩展、可靠、有效的反馈路径。

原文摘要 · Abstract (English)

Objective and scalable measurement of teaching quality is a persistent challenge in education. While Large Language Models (LLMs) offer potential, general-purpose models have struggled to reliably apply complex, authentic classroom observation instruments. This paper uses custom LLMs built on sentence-level embeddings, an architecture better suited for the long-form, interpretive nature of classroom transcripts than conventional subword tokenization. We systematically evaluate five different sentence embeddings under a data-efficient training regime designed to prevent overfitting. Our results demonstrate that these specialized models can achieve human-level and even super-human performance with expert human ratings above 0.65 and surpassing the average human-human rater correlation. Further, through analysis of annotation context windows, we find that more advanced models-those better aligned with human judgments-attribute a larger share of score variation to lesson-level features rather than isolated utterances, challenging the sufficiency of single-turn annotation paradigms. Finally, to assess external validity, we find that aggregate model scores align with teacher value-added measures, indicating they are capturing features relevant to student learning. However, this trend does not hold at the individual item level, suggesting that while the models learn useful signals, they have not yet achieved full generalization. This work establishes a viable and powerful new methodology for AI-driven instructional measurement, offering a path toward providing scalable, reliable, and valid feedback for educator development.

教学评估大模型应用教育科技

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。