用多智能体对话评估大模型教学能力,发现小模型有时比大模型更会教。
EducationQ: Evaluating LLMs' Teaching Capabilities Through Multi-Agent Dialogue Framework
- 设计教学、学习、评估三类智能体,模拟真实课堂互动。
- 14个模型在13个学科10个难度上测试,小模型表现超大模型。
- 强调教学策略和反馈适应性,适合教育AI研发者参考。
大型语言模型(LLMs)日益成为教育工具,但因其资源密集、情境依赖和方法复杂,评估其教学能力仍具挑战。本文提出EducationQ,一种基于多智能体对话的框架,通过模拟动态教育场景高效评估教学能力,包含教学、学习与评估三类专用智能体。在涵盖13个学科、10个难度层级的1,498道题上测试14个来自OpenAI、Meta、Google、Anthropic等机构的LLMs,结果表明教学效果与模型规模或通用推理能力无线性关联——部分小型开源模型在教学任务中优于大型商用模型。该发现揭示了现有评估体系偏重知识回忆而忽视交互式教学的显著缺陷。混合方法评估结合定量指标、定性分析与专家案例研究,识别出顶尖模型的差异化教学优势(如高阶提问策略、自适应反馈机制)。人类专家对有效教学行为的判断与自动化分析达成78%一致性,验证了方法可靠性。EducationQ表明,将LLM作为教师需超越单纯规模扩展,进行针对性教学能力优化,提示下一代教育AI应聚焦特定教学效能提升。
原文摘要 · Abstract (English)
Large language models (LLMs) increasingly serve as educational tools, yet evaluating their teaching capabilities remains challenging due to the resource-intensive, context-dependent, and methodologically complex nature of teacher-student interactions. We introduce EducationQ, a multi-agent dialogue framework that efficiently assesses teaching capabilities through simulated dynamic educational scenarios, featuring specialized agents for teaching, learning, and evaluation. Testing 14 LLMs across major AI Organizations (OpenAI, Meta, Google, Anthropic, and others) on 1,498 questions spanning 13 disciplines and 10 difficulty levels reveals that teaching effectiveness does not correlate linearly with model scale or general reasoning capabilities - with some smaller open-source models outperforming larger commercial counterparts in teaching contexts. This finding highlights a critical gap in current evaluations that prioritize knowledge recall over interactive pedagogy. Our mixed-methods evaluation, combining quantitative metrics with qualitative analysis and expert case studies, identifies distinct pedagogical strengths employed by top-performing models (e.g., sophisticated questioning strategies, adaptive feedback mechanisms). Human expert evaluations show 78% agreement with our automated qualitative analysis of effective teaching behaviors, validating our methodology. EducationQ demonstrates that LLMs-as-teachers require specialized optimization beyond simple scaling, suggesting next-generation educational AI prioritize targeted enhancement of specific pedagogical effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。