评测大模型能否像真人导师一样根据学生状态动态调整教学策略。
Discerning minds or generic tutors? Evaluating instructional guidance capabilities in Socratic LLMs
- 构建三阶段评估框架,分析模型如何感知、调整和引导学习者。
- 现有模型在学生困惑时缺乏有效引导,表现远低于专家水平。
- 提出行为引导微调方法,显著提升教学适应性,适合教育AI研究者。
大型语言模型的对话能力为可扩展、互动式辅导带来巨大潜力。以往研究多关注生成苏格拉底式问题,却忽视了关键环节:根据学习者的认知状态进行自适应指导。本研究超越问题生成,聚焦教学指导能力。核心问题是:大模型能否像专家导师一样,依据学习者状态动态调整教学策略?为此,我们提出GuideEval基准,基于真实教育对话,通过三阶段行为框架评估教学指导能力:(1)感知,推断学习者状态;(2)协调,调整教学策略;(3)激发,引导适当反思。实证结果表明,现有大模型在学习者困惑或需引导时,往往无法提供有效的自适应支架。为补充量化评估,我们进行了细致的失败案例分析,直观揭示其局限。此外,我们引入行为引导微调策略,利用行为提示的教学对话,显著提升指导性能。本研究推动评价范式从孤立内容评估转向以学习者为中心的状态感知交互,倡导更对话化的苏格拉底式大模型评估路径。
原文摘要 · Abstract (English)
The conversational capabilities of large language models hold significant promise for enabling scalable and interactive tutoring. While prior research has primarily examined their ability to generate Socratic questions, it often overlooks a critical aspect: adaptively guiding learners in accordance with their cognitive states. This study moves beyond question generation to emphasize instructional guidance capability. We ask: Can LLMs emulate expert tutors who dynamically adjust strategies in response to learners' states? To investigate this, we propose GuideEval, a benchmark grounded in authentic educational dialogues that evaluates pedagogical guidance through a three-phase behavioral framework: (1) Perception, inferring learner states; (2) Orchestration, adapting instructional strategies; and (3) Elicitation, stimulating proper reflections. Empirical results indicate that existing LLMs often fail to provide effective adaptive scaffolding when learners experience confusion or require redirection. To complement the quantitative evaluation, we conduct a detailed failure case analysis, providing an intuitive understanding of these shortcomings. Furthermore, we introduce a behavior-guided finetuning strategy that leverages behavior-prompted instructional dialogues, substantially enhancing guidance performance. By shifting the focus from isolated content evaluation to learner-centered state-aware interaction, our work advocates a more dialogic paradigm for evaluating Socratic LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。