arXiv:2411.03495cs.CLcs.AI2024-11中稿 · NeurIPS被引 15

用大模型自动生成数学题提示,帮助学生纠正错误。

Automatic Generation of Question Hints for Mathematics Problems using Large Language Models in Educational Technology

  • 用大模型模拟学生做题,分析错因并生成针对性提示。
  • 针对特定错误的提示比通用提示更有效,低温度设置效果更好。
  • 开源模型作为教师表现优于部分闭源模型,适合教育应用。

大型语言模型(LLMs)在智能辅导系统中自动生成提示具有提升学习效果的潜力。然而,如何生成符合教学原则、能纠正学生误解并达成教育目标的提示仍具挑战。本研究利用GPT-4o和Llama-3-8B-instruct作为教师,通过GPT-3.5-turbo、Llama-3-8B-Instruct或Mistral-7B-instruct-v0.3模拟高中生解题,基于认知科学设计数学题目。研究涵盖三方面:1)识别模拟学生在中学数学题中的错误模式;2)设计多种提示给GPT-4o生成提示,并评估其帮助学生自我修正的能力;3)将表现最佳的提示用于Llama-3-8B-Instruct作为教师,与GPT-4o进行性能对比。结果表明,模型错误随温度升高而增加。当由GPT-4o生成提示时,针对具体错误的提示及基于常见错误的通用提示最有效。有趣的是,使用Llama-3-8B-Instruct作为教师整体表现优于GPT-4o。此外,学生模型(尤其是GPT-3.5-turbo)在获得提示后,解题与修改能力显著提升,尤其在低温度设置下;而Mistral-7B-Instruct在高温下性能下降。

原文摘要 · Abstract (English)

The automatic generation of hints by Large Language Models (LLMs) within Intelligent Tutoring Systems (ITSs) has shown potential to enhance student learning. However, generating pedagogically sound hints that address student misconceptions and adhere to specific educational objectives remains challenging. This work explores using LLMs (GPT-4o and Llama-3-8B-instruct) as teachers to generate effective hints for students simulated through LLMs (GPT-3.5-turbo, Llama-3-8B-Instruct, or Mistral-7B-instruct-v0.3) tackling math exercises designed for human high-school students, and designed using cognitive science principles. We present here the study of several dimensions: 1) identifying error patterns made by simulated students on secondary-level math exercises; 2) developing various prompts for GPT-4o as a teacher and evaluating their effectiveness in generating hints that enable simulated students to self-correct; and 3) testing the best-performing prompts, based on their ability to produce relevant hints and facilitate error correction, with Llama-3-8B-Instruct as the teacher, allowing for a performance comparison with GPT-4o. The results show that model errors increase with higher temperature settings. Notably, when hints are generated by GPT-4o, the most effective prompts include prompts tailored to specific errors as well as prompts providing general hints based on common mathematical errors. Interestingly, Llama-3-8B-Instruct as a teacher showed better overall performance than GPT-4o. Also the problem-solving and response revision capabilities of the LLMs as students, particularly GPT-3.5-turbo, improved significantly after receiving hints, especially at lower temperature settings. However, models like Mistral-7B-Instruct demonstrated a decline in performance as the temperature increased.

教育AI大模型数学题提示生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。