首个针对AI教学适配能力的评测基准,揭示当前模型内容强但因材施教弱。
Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents

- 以学习者画像为评估标准,生成完整教学视频
- 顶尖模型在适应性上表现相近,自动评分无法区分优劣
- 适用于教育AI、评测基准研究者
AI代理如今能解题、像专家作答并生成长篇多模态内容,但能否根据特定学习者调整教学,即教育学中的教学内容知识(PCK),尚未有评测标准。为此,我们推出「教学怪物挑战赛」,首个将学习者画像作为显式评估指标的教学视频生成基准。每个系统接收一个主题和学习者画像,需生成完整教学视频。视频由LLM裁判筛选、人群两两投票排序,最终由专家小组定稿。首期结果表明,当前系统内容掌握良好,但在呈现与适配方面明显不足;同时暴露自动评判的局限:LLM裁判虽能识别低分尾部,但对顶尖系统评分趋同,其排名与人类偏好不一致。因此,进步不仅需要更优教学系统,还需更好自动评判机制。我们发布该基准、评分细则及人工标注数据,作为双方面测试平台。
原文摘要 · Abstract (English)
AI agents can now solve problems, answer like subject experts, and generate long-form multimodal content. However, whether they can adapt a lesson to fit a specified learner, which education calls Pedagogical Content Knowledge (PCK), has not been benchmarked. To measure it, we introduce the Teaching Monster Challenge, the first instructional video generation benchmark to treat the learner persona as an explicit evaluation criterion. Each system is given a topic and a learner persona and must generate a complete instructional video. Every video is screened by an LLM-judge, ranked by crowd pairwise voting, and finalized by an expert panel. The first edition shows that today's systems handle the content well but are far weaker at presenting it and adapting it to the learner. The same process exposes a limit of automatic judging. The LLM-judge separates a clear low-performing tail but ranks the strongest systems poorly. The strongest systems receive nearly identical scores from the judge, so its ranking of them does not match human preference. Progress therefore requires not only better teaching systems but also better automatic judges, and we release the benchmark, rubric, and human judgments as a testbed for both.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。