arXiv:2601.21375cs.AI2026-01

用课程大纲评估大模型教学能力,发现数学教得好,物理化学却难。

TeachBench: A Syllabus-Grounded Framework for Evaluating Teaching Ability in Large Language Models

  • 基于课程大纲设计多轮教学评估,防止知识泄露
  • 数学模型教学效果好,物化类科目表现差
  • 示例题未必提升教学,可能诱导死记硬背

大型语言模型(LLMs)在作为教学助手方面展现出潜力,但其教学能力尚未得到充分评估。现有基准主要关注解题或问题层面的指导,对以知识为中心的教学关注不足。本文提出一种基于课程大纲的评估框架,通过学生在多轮教学后的表现提升来衡量大模型的教学能力。通过限制教师代理仅使用结构化知识点和例题,该框架避免信息泄露,并可复用现有基准。我们在高考数据集上对多个学科进行了实例化。实验显示不同模型和学科间教学效果差异显著:部分模型在数学上表现良好,但在物理和化学领域仍具挑战性。此外,引入例题并不必然提升教学效果,因为模型常转向针对例题的错误修正。总体而言,研究结果凸显教学能力是大模型行为中一个独特且可度量的维度。

原文摘要 · Abstract (English)

Large language models (LLMs) show promise as teaching assistants, yet their teaching capability remains insufficiently evaluated. Existing benchmarks mainly focus on problem-solving or problem-level guidance, leaving knowledge-centered teaching underexplored. We propose a syllabus-grounded evaluation framework that measures LLM teaching capability via student performance improvement after multi-turn instruction. By restricting teacher agents to structured knowledge points and example problems, the framework avoids information leakage and enables reuse of existing benchmarks. We instantiate the framework on Gaokao data across multiple subjects. Experiments reveal substantial variation in teaching effectiveness across models and domains: some models perform well in mathematics, while teaching remains challenging in physics and chemistry. We also find that incorporating example problems does not necessarily improve teaching, as models often shift toward example-specific error correction. Overall, our results highlight teaching ability as a distinct and measurable dimension of LLM behavior.

教学评估大模型课程大纲考试数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。