用强化学习训练AI导师,让对话中学生答对率更高
Training LLM-based Tutors to Improve Student Learning Outcomes in Dialogues
- 用学生模型和教学评分双重打分,筛选最优回应
- 训练后学生答对率显著提升,教学品质不下降
- 适合教育科技、智能辅导系统研发者参考
生成式人工智能有潜力通过大语言模型(LLMs)实现个性化辅导的规模化。尽管近期的AI导师通过微调或提示使模型遵循有效教学原则,但尚未针对对话全程中的学生学习效果进行优化,可能导致教学方式不够高效。为此,我们提出一种新方法:训练LLM生成能最大化学生答对概率的导师回应,同时保持良好的教学实践。具体而言,我们生成一组候选回应,分别通过(1)基于LLM的学生模型预测学生正确回答的概率,以及(2)由GPT-4o评估的教学质量评分标准进行打分。利用这些数据,我们采用直接偏好优化(DPO)训练开源模型Llama 3.1 8B。实验表明,本模型生成的导师回应显著提高学生答对概率,且教学品质与GPT-4o相当。定性分析与人工评估也证实了生成内容的质量。
原文摘要 · Abstract (English)
Generative artificial intelligence (AI) has the potential to scale up personalized tutoring through large language models (LLMs). Recent AI tutors are adapted for the tutoring task by training or prompting LLMs to follow effective pedagogical principles, though they are not trained to maximize student learning throughout the course of a dialogue. Therefore, they may engage with students in a suboptimal way. We address this limitation by introducing an approach to train LLMs to generate tutor utterances that maximize the likelihood of student correctness, while still encouraging the model to follow good pedagogical practice. Specifically, we generate a set of candidate tutor utterances and score them using (1) an LLM-based student model to predict the chance of correct student responses and (2) a pedagogical rubric evaluated by GPT-4o. We then use the resulting data to train an open-source LLM, Llama 3.1 8B, using direct preference optimization. We show that tutor utterances generated by our model lead to significantly higher chances of correct student responses while maintaining the pedagogical quality of GPT-4o. We also conduct qualitative analyses and a human evaluation to demonstrate that our model generates high quality tutor utterances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。