构建了覆盖345个知识点的师生对话数据集,用于评估大模型教学能力。
EduDial: Constructing a Large-scale Multi-turn Teacher-Student Dialogue Corpus
- 基于布鲁姆教育目标分类设计多轮师生对话,融合十种提问策略。
- 涵盖34,250个对话会话,覆盖345个核心知识点,支持个性化教学评估。
- 适用于教育AI研究者与开发者,尤其关注智能辅导系统评测。
近年来,多个多轮对话基准被提出以评估大语言模型(LLMs)的对话能力。随着LLMs在智能教育领域的重要性日益凸显,因其能深入理解教学情境并提供个性化指导,专门的师生对话基准建设变得尤为关键。为此,我们提出EduDial,一个全面的多轮师生对话数据集。EduDial涵盖345个核心知识点,包含34,250个由教师与学生代理交互生成的对话会话。其设计基于布鲁姆教育目标分类法,并融入十种提问策略,包括情境提问、最近发展区(ZPD)提问和元认知提问,更真实地捕捉课堂互动。此外,针对不同认知水平的学生设计差异化教学策略,实现更精准的教学引导。基于EduDial,我们通过训练构建了EduDial-LLM 32B,并提出一个11维评估框架,系统性衡量LLMs的教学能力,涵盖整体教学质量与内容质量。对17个主流LLMs的实验表明,多数模型在以学生为中心的教学场景中表现不佳,而我们的EduDial-LLM在所有指标上均显著优于基线。代码已开源。
原文摘要 · Abstract (English)
Recently, several multi-turn dialogue benchmarks have been proposed to evaluate the conversational abilities of large language models (LLMs). As LLMs are increasingly recognized as a key technology for advancing intelligent education, owing to their ability to deeply understand instructional contexts and provide personalized guidance, the construction of dedicated teacher-student dialogue benchmarks has become particularly important. To this end, we present EduDial, a comprehensive multi-turn teacher-student dialogue dataset. EduDial covers 345 core knowledge points and consists of 34,250 dialogue sessions generated through interactions between teacher and student agents. Its design is guided by Bloom's taxonomy of educational objectives and incorporates ten questioning strategies, including situational questioning, zone of proximal development (ZPD) questioning, and metacognitive questioning-thus better capturing authentic classroom interactions. Furthermore, we design differentiated teaching strategies for students at different cognitive levels, thereby providing more targeted teaching guidance. Building on EduDial, we further develop EduDial-LLM 32B via training and propose an 11-dimensional evaluation framework that systematically measures the teaching abilities of LLMs, encompassing both overall teaching quality and content quality. Experiments on 17 mainstream LLMs reveal that most models struggle in student-centered teaching scenarios, whereas our EduDial-LLM achieves significant gains, consistently outperforming all baselines across all metrics. The code is available at https://github.com/Mind-Lab-ECNU/EduDial/tree/main.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。