用多智能体模拟教学对话,实现规模化程序化学习与教学质量评估
Conversational Education at Scale: A Multi-LLM Agent Workflow for Procedural Learning and Pedagogic Quality Assessment
- 构建教师-学习者-管理器-评估器多智能体流程,模拟真实教学互动
- 基于14,287篇教程生成11.4万组对话,覆盖17个领域727个主题
- 开源数据集与代码,适合教育科技、AI教学研究者使用
大型语言模型(LLMs)已推动虚拟教与学的发展,融合自然语言处理与AI for Education。现有研究常缺乏可扩展性,难以利用多样化大规模课程内容,且缺少教学品质评估框架。为此,我们提出WikiHowAgent,一种基于LLM的多智能体工作流,用于模拟交互式教学-学习对话。该系统整合教师与学习者智能体、交互管理器及评估器,支持程序化学习并评估教学品质。我们构建了一个包含114,296组师生对话的数据集,源自14,287篇教程,覆盖17个领域与727个主题。评估协议结合计算指标与评分标准,辅以人工判断对齐。实验表明该流程在多种场景下均具有效性,揭示了LLM在不同领域的能力表现。所有数据集与实现均已开源。
原文摘要 · Abstract (English)
Large language models (LLMs) have advanced virtual educators and learners, bridging NLP with AI4Education. Existing work often lacks scalability and fails to leverage diverse, large-scale course content, with limited frameworks for assessing pedagogic quality. To this end, we propose WikiHowAgent, a multi-agent workflow leveraging LLMs to simulate interactive teaching-learning conversations. It integrates teacher and learner agents, an interaction manager, and an evaluator to facilitate procedural learning and assess pedagogic quality. We introduce a dataset of 114,296 teacher-learner conversations grounded in 14,287 tutorials across 17 domains and 727 topics. Our evaluation protocol combines computational and rubric-based metrics with human judgment alignment. Results demonstrate the workflow's effectiveness in diverse setups, offering insights into LLM capabilities across domains. Our datasets and implementations are fully open-sourced.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。