首个系统评估AI生成教案的基准数据集,覆盖240个理科主题。
LessonBench-V1: A Benchmark Dataset for Evaluating AI Lesson Generation Agents

- 基于教育学理论构建240个跨学科教案计划,含3620个学习目标。
- 整合97个权威开源资源,每份教案配有人工审核的反向工程计划。
- 提供三维评估框架,支持可复现的AI教案生成效果评测。
基于大语言模型的AI教育内容生成系统日益发展,但缺乏标准化评估基准。本文提出LessonBench-V1,一个包含647份人工撰写教案及其对应LLM反向生成教案计划的数据集,覆盖数学、物理、化学和计算机科学共240个STEM主题。教案源自LibreTexts、Brilliant.org、GeeksForGeeks等97个可信开放源。每个教案计划均经人工评审,采用结合布卢姆分类法、加涅五阶段、梅里尔首要原则及5E教学模型的教育学方法论生成。数据集共包含3,620个学习目标及教学元数据,支持对教案生成类AI代理进行系统化、可复现的评估,并提出配套的三维度评估流程,推动该领域研究发展。
原文摘要 · Abstract (English)
Large Language Model (LLM) based AI educational content generation systems are increasingly being developed, yet no standardised benchmark exists to systematically evaluate them. This study introduces LessonBench-V1, a benchmark dataset comprising 647 human-written lessons paired with LLM-based reverse-engineered lesson plans across 240 STEM topics spanning mathematics, physics, chemistry, and computer science. The lessons are drawn from 97 trusted open sources, including LibreTexts, Brilliant.org and GeeksForGeeks. Each lesson plan is human-reviewed and produced through a pedagogically grounded methodology that synthesises Bloom's Taxonomy, Gagné's Events, Merrill's First Principles, and the 5E Instructional Model. The lesson plans capture 3,620 learning objectives with pedagogical metadata, enabling systematic, reproducible evaluation of lesson-generation AI agents and supporting further research. The study further proposes a three-dimensional evaluation pipeline for use with the dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。