arXiv:2507.20335cs.LGcs.AI2025-07被引 1

用强化学习让AI导师更懂教学,更贴心更有创意。

Cultivating Helpful, Personalized, and Creative AI Tutors: A Framework for Pedagogical Alignment using Reinforcement Learning

  • 构建三维度评分体系,用奖励模型评估AI教学表现。
  • 在8000条教育对话上训练,2000个提示微调后效果显著提升。
  • 适合想打造智能、个性化、有创造力的教育AI的研究者。

将大语言模型(LLMs)融入教育,可实现大规模个性化学习,但通用模型常缺乏教学契合度,如帮助性、学生中心化和创造力培养。为此,本文提出EduAlign框架,分两阶段实现教育对齐。第一阶段,构建包含8000条教育交互的数据集,从帮助性、个性化、创造力(HPC)三个维度进行人工与自动标注,并训练出多维奖励模型HPC-RM,用于精准评分。第二阶段,使用HPC-RM作为奖励信号,基于2000个多样化提示,通过组相对策略优化(GRPO)微调预训练模型。实验表明,微调后模型在教育与通用基准上均显著提升对三维度的对齐程度。该方法为打造更具参与感、符合教学规律的AI导师提供了可扩展的有效路径。

原文摘要 · Abstract (English)

The integration of large language models (LLMs) into education presents unprecedented opportunities for scalable personalized learning. However, standard LLMs often function as generic information providers, lacking alignment with fundamental pedagogical principles such as helpfulness, student-centered personalization, and creativity cultivation. To bridge this gap, we propose EduAlign, a novel framework designed to guide LLMs toward becoming more effective and responsible educational assistants. EduAlign consists of two main stages. In the first stage, we curate a dataset of 8k educational interactions and annotate them-both manually and automatically-along three key educational dimensions: Helpfulness, Personalization, and Creativity (HPC). These annotations are used to train HPC-RM, a multi-dimensional reward model capable of accurately scoring LLM outputs according to these educational principles. We further evaluate the consistency and reliability of this reward model. In the second stage, we leverage HPC-RM as a reward signal to fine-tune a pre-trained LLM using Group Relative Policy Optimization (GRPO) on a set of 2k diverse prompts. We then assess the pre- and post-finetuning models on both educational and general-domain benchmarks across the three HPC dimensions. Experimental results demonstrate that the fine-tuned model exhibits significantly improved alignment with pedagogical helpfulness, personalization, and creativity stimulation. This study presents a scalable and effective approach to aligning LLMs with nuanced and desirable educational traits, paving the way for the development of more engaging, pedagogically aligned AI tutors.

AI教育强化学习个性化教学语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。