测试大模型对教学方法和特殊教育的理解能力,发现差距显著。
Benchmarking the Pedagogical Knowledge of Large Language Models
- 用教师培训考试题构建新基准,评估跨领域教学知识
- 97个模型准确率28%至89%,表现差异大
- 适合教育AI研发者与政策制定者参考
MMLU等基准在评估AI多领域知识方面作用关键,但主要关注内容知识,忽视了对教学方法(pedagogy)的评测。本文提出《教学能力基准》(The Pedagogy Benchmark),用于评估大语言模型在跨领域教学知识(CDPK)和特殊教育需求与残疾(SEND)教学知识方面的能力。该基准基于教师专业发展考试中的精选题目,覆盖教学策略、评估方法等子领域。我们报告了97个模型在教学知识题上的表现,准确率范围为28%至89%。分析成本与准确率关系,绘制帕累托前沿随时间变化趋势。提供在线排行榜(https://rebrand.ly/pedagogy),支持按每令牌成本、模型开源状态及不同学科性能进行交互筛选。教育类大模型潜力巨大,需此类基准来衡量其理解教学概念、回应学习者需求、支持有效教学实践的能力,以指导负责任且基于证据的部署与政策制定。
原文摘要 · Abstract (English)
Benchmarks like Massive Multitask Language Understanding (MMLU) have played a pivotal role in evaluating AI's knowledge and abilities across diverse domains. However, existing benchmarks predominantly focus on content knowledge, leaving a critical gap in assessing models' understanding of pedagogy - the method and practice of teaching. This paper introduces The Pedagogy Benchmark, a novel dataset designed to evaluate large language models on their Cross-Domain Pedagogical Knowledge (CDPK) and Special Education Needs and Disability (SEND) pedagogical knowledge. These benchmarks are built on a carefully curated set of questions sourced from professional development exams for teachers, which cover a range of pedagogical subdomains such as teaching strategies and assessment methods. Here we outline the methodology and development of these benchmarks. We report results for 97 models, with accuracies spanning a range from 28% to 89% on the pedagogical knowledge questions. We consider the relationship between cost and accuracy and chart the progression of the Pareto value frontier over time. We provide online leaderboards at https://rebrand.ly/pedagogy which are updated with new models and allow interactive exploration and filtering based on various model properties, such as cost per token and open-vs-closed weights, as well as looking at performance in different subjects. LLMs and generative AI have tremendous potential to influence education and help to address the global learning crisis. Education-focused benchmarks are crucial to measure models' capacities to understand pedagogical concepts, respond appropriately to learners' needs, and support effective teaching practices across diverse contexts. They are needed for informing the responsible and evidence-based deployment of LLMs and LLM-based tools in educational settings, and for guiding both development and policy decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。