用真实师生对话数据微调模型,让AI教学更像真人。
TeachLM: Post-Training LLMs for Education Using Authentic Learning Data
- 基于10万小时真实师生对话,通过参数高效微调构建教学模型。
- 微调后对话轮次增50%,学生发言时间翻倍,提问更自然。
- 适合教育AI研发者、教学工具设计师,提升对话教学能力。
生成式AI有望重塑教育,但大语言模型的教育能力受限于缺乏真实学习数据。当前依赖提示工程作为临时方案,但其在编码复杂教学策略方面存在根本局限。为此,我们提出TeachLM——一种通过参数高效微调优化的教学型大模型。该模型基于由Polygence维护的10万小时一对一、纵向师生互动数据集训练,经过严格匿名化处理以保护隐私。我们利用该模型生成高保真度的合成师生对话,并提出一种新型多轮评估协议,实现快速、可扩展且可复现的对话能力评测。实验表明,使用真实学习数据微调显著提升对话与教学表现:学生发言时间翻倍,提问风格改善,对话轮次增加50%,教学个性化程度更高。
原文摘要 · Abstract (English)
The promise of generative AI to revolutionize education is constrained by the pedagogical limits of large language models (LLMs). A major issue is the lack of access to high-quality training data that reflect the learning of actual students. Prompt engineering has emerged as a stopgap, but the ability of prompts to encode complex pedagogical strategies in rule-based natural language is inherently limited. To address this gap we introduce TeachLM - an LLM optimized for teaching through parameter-efficient fine-tuning of state-of-the-art models. TeachLM is trained on a dataset comprised of 100,000 hours of one-on-one, longitudinal student-tutor interactions maintained by Polygence, which underwent a rigorous anonymization process to protect privacy. We use parameter-efficient fine-tuning to develop an authentic student model that enables the generation of high-fidelity synthetic student-tutor dialogues. Building on this capability, we propose a novel multi-turn evaluation protocol that leverages synthetic dialogue generation to provide fast, scalable, and reproducible assessments of the dialogical capabilities of LLMs. Our evaluations demonstrate that fine-tuning on authentic learning data significantly improves conversational and pedagogical performance - doubling student talk time, improving questioning style, increasing dialogue turns by 50%, and greater personalization of instruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。