首个评估大模型跨学科引导能力的基准,助力教育AI发展
SID: Benchmarking Guided Instruction Capabilities in STEM Education with a Socratic Interdisciplinary Dialogues Dataset
- 构建1万次对话的跨学科苏格拉底式对话数据集
- 发现顶尖大模型仍难以有效引导知识整合与迁移
- 适合教育科技、AI教学助手研发者参考
培养学生在复杂问题求解中实现知识整合与迁移的能力是现代教育的核心目标,跨学科STEM教育是关键路径,但需要专家指导且难以规模化。尽管大语言模型(LLMs)在此领域具有潜力,但其真实引导能力因缺乏有效评估基准而不明。为此,我们提出SID,首个系统评估LLMs在多轮跨学科苏格拉底对话中高阶引导能力的基准。包含48个复杂STEM项目、共10,000次对话轮次的大规模数据集,创新的标注框架捕捉深层教学特征,并引入新评估指标(如X-SRG)。基线实验表明,即使最先进的大模型也难以开展有效引导对话以促成知识整合与迁移。该基准凸显了推动更具教育意识的LLMs发展的关键价值。
原文摘要 · Abstract (English)
Fostering students' abilities for knowledge integration and transfer in complex problem-solving scenarios is a core objective of modern education, and interdisciplinary STEM is a key pathway to achieve this, yet it requires expert guidance that is difficult to scale. While LLMs offer potential in this regard, their true capability for guided instruction remains unclear due to the lack of an effective evaluation benchmark. To address this, we introduce SID, the first benchmark designed to systematically evaluate the higher-order guidance capabilities of LLMs in multi-turn, interdisciplinary Socratic dialogues. Our contributions include a large-scale dataset of 10,000 dialogue turns across 48 complex STEM projects, a novel annotation schema for capturing deep pedagogical features, and a new suite of evaluation metrics (e.g., X-SRG). Baseline experiments confirm that even state-of-the-art LLMs struggle to execute effective guided dialogues that lead students to achieve knowledge integration and transfer. This highlights the critical value of our benchmark in driving the development of more pedagogically-aware LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。