用合成数据微调小模型,实现低成本高效一对一辅导。
Developing a Tutoring Dialog Dataset to Optimize LLMs for Educational Use
- 用合成对话数据训练小型LLM,替代昂贵专家数据
- 微调后的小模型表现媲美大模型,成本更低
- 适合预算有限的教育机构部署智能辅导系统
大型语言模型在教育应用中展现出潜力,但基于对话的辅导系统仍面临教学策略有效性和专家标注数据高成本的挑战。本研究探索使用更小、更经济的LLM在阅读理解问题的一对一辅导场景中的应用。我们构建了一个合成的辅导对话数据集,并由人类教师评估其质量;随后使用该数据集微调一个小型LLM。此外,我们在真实场景中开展交互实验,比较微调后的模型与大型模型的表现。结果表明,微调后的模型性能与大模型相当,但成本显著降低,验证了在教育环境中实现低成本、可扩展的LLM辅导系统的可行性。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have shown promise for scalable educational applications, but their use in dialog-based tutoring systems remains challenging due to the need for effective pedagogical strategies and the high costs associated with expert-curated datasets. Our study explores the use of smaller, more affordable LLMs for one-on-one tutoring in the context of solving reading comprehension problems. We developed a synthetic tutoring dialog dataset, evaluated by human teachers, and fine-tuned a smaller LLM using this dataset. Furthermore, we conducted an interactive experiment comparing the performance of the fine-tuned model with a larger model in real-world tutoring scenarios. Our results show that the fine-tuned model performs on par with the larger model but at a lower cost, demonstrating a viable, cost-effective approach for implementing LLM-based tutoring systems in educational settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。