arXiv:2602.10620cs.SEcs.CL2026-02

首个系统评估LLM教学设计智能体的基准,验证理论结合推理效果最佳。

ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents

  • 构建25,795个场景的基准,融合51个变量与ADDIE模型33步流程
  • 理论+推理的智能体在1,017个测试中表现最优,显著优于纯技术或纯理论方案
  • 适合教育科技研发者、AI教育应用开发者参考

大型语言模型(LLM)代理在自动化教学系统设计(ISD)方面展现出巨大潜力,但评估仍面临基准缺失和模型自评偏差问题。本文提出ISD-Agent-Bench,一个涵盖25,795个场景的综合性基准,基于上下文矩阵框架,整合51个跨5类情境变量与来自ADDIE模型的33个子步骤。为保障评估可靠性,采用多模型评判协议,使用不同厂商的多种LLM,实现高评委间一致性。在1,017个测试场景中,对比现有及基于经典ISD理论(如ADDIE、Dick & Carey、快速原型法)的新代理,发现将传统理论与现代ReAct式推理结合的方案性能最高,显著优于纯理论或纯技术方法。进一步分析表明,理论质量与基准表现强相关,理论型代理在以问题为中心设计和目标-评估对齐上具有显著优势。本工作为系统性开展基于LLM的ISD研究奠定基础。

原文摘要 · Abstract (English)

Large Language Model (LLM) agents have shown promising potential in automating Instructional Systems Design (ISD), a systematic approach to developing educational programs. However, evaluating these agents remains challenging due to the lack of standardized benchmarks and the risk of LLM-as-judge bias. We present ISD-Agent-Bench, a comprehensive benchmark comprising 25,795 scenarios generated via a Context Matrix framework that combines 51 contextual variables across 5 categories with 33 ISD sub-steps derived from the ADDIE model. To ensure evaluation reliability, we employ a multi-judge protocol using diverse LLMs from different providers, achieving high inter-judge reliability. We compare existing ISD agents with novel agents grounded in classical ISD theories such as ADDIE, Dick \& Carey, and Rapid Prototyping ISD. Experiments on 1,017 test scenarios demonstrate that integrating classical ISD frameworks with modern ReAct-style reasoning achieves the highest performance, outperforming both pure theory-based agents and technique-only approaches. Further analysis reveals that theoretical quality strongly correlates with benchmark performance, with theory-based agents showing significant advantages in problem-centered design and objective-assessment alignment. Our work provides a foundation for systematic LLM-based ISD research.

教学设计LLM评估智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。