测试大模型在真实课堂中识别教学行为的能力,发现提示设计能显著提升表现。
How well do Large Language Models Recognize Instructional Moves? Establishing Baselines for Foundation Models in Educational Discourse
- 用零样本、单样本和少样本提示对比六种大模型对课堂对话的教学行为分类能力。
- 最佳配置下模型达成0.58的克里特·卡帕系数,但不同教学行为表现差异大。
- 适合教育领域研究者和智能教学系统开发者参考模型的实用边界。
大型语言模型(LLMs)在教育技术中被广泛用于生成教学材料、辅助评估设计和辅导等任务。然而,关于这些模型在未经显著定制的情况下理解真实教育场景的能力仍知之甚少。随着基于LLM的系统在日常教学中日益普及,了解其开箱即用的表现至关重要。我们比较了六种LLM在一项基础但关键的任务上的表现:对真实课堂转录文本中的教学行为进行分类。评估了典型的提示方法——零样本、单样本和少样本提示。结果显示,尽管零样本性能中等,但提供详尽示例(少样本提示)显著提升了先进模型的表现,最强配置达到与专家标注一致的Cohen's Kappa = 0.58。同时,提升并非均匀或完整:不同教学行为的表现差异显著,高召回率常伴随更高的误报率。总体而言,这些发现表明基础模型在解读教学话语方面具有有意义但有限的能力,提示设计有助于释放潜力,却无法消除根本的可靠性限制。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly adopted in educational technologies for a variety of tasks, from generating instructional materials and assisting with assessment design to tutoring. While prior work has investigated how models can be adapted or optimized for specific tasks, far less is known about how well LLMs perform at interpreting authentic educational scenarios without significant customization. As LLM-based systems become widely adopted by learners and educators in everyday academic contexts, understanding their out-of-the-box capabilities is increasingly important for setting expectations and benchmarking. We compared six LLMs to estimate their baseline performance on a simple but important task: classifying instructional moves in authentic classroom transcripts. We evaluated typical prompting methods: zero-shot, one-shot, and few-shot prompting. We found that while zero-shot performance was moderate, providing comprehensive examples (few-shot prompting) significantly improved performance for state-of-the-art models, with the strongest configuration reaching Cohen's Kappa = 0.58 against expert-coded annotations. At the same time, improvements were neither uniform nor complete: performance varied considerably by instructional move, and higher recall frequently came at the cost of increased false positives. Overall, these findings indicate that foundation models demonstrate meaningful yet limited capacity to interpret instructional discourse, with prompt design helping to surface capability but not eliminating fundamental reliability constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。