用苏格拉底式对话+游戏化,让大模型更懂教数学。
Findings of MEGA: Maths Explanation with LLMs using the Socratic Method for Active Learning
- 用苏格拉底提问+游戏化反馈,引导学生主动思考数学。
- 在MATH数据集上,MEGA方法被认可度达47.5%,远超传统CoT的26.67%。
- 适合想提升数学理解力的学生或教育科技研究者。
本文开展干预实验,研究苏格拉底教学法、思维链(CoT)推理、简化游戏化及形成性反馈结合对大学生数学学习的影响。提出基于AI大模型的数学解释系统MEGA。针对学生普遍存在的数学畏难情绪及其根源——教学方法不当,采用随机分组设计,对比了MEGA与传统步骤式(CoT)方法。从GSM8K和MATH两个数据集中分别随机抽取样本(n=60),误差范围11%,置信水平90%,用于评估六种候选大模型中表现最优的两种:GPT4o与Claude 3.5 Sonnet。结果显示,多数学生认为MEGA在两个数据集上都更利于学习,尤其在较难的MATH数据集上,认可度达47.5%,显著高于传统CoT的26.67%。
原文摘要 · Abstract (English)
This paper presents an intervention study on the effects of the combined methods of (1) the Socratic method, (2) Chain of Thought (CoT) reasoning, (3) simplified gamification and (4) formative feedback on university students' Maths learning driven by large language models (LLMs). We call our approach Mathematics Explanations through Games by AI LLMs (MEGA). Some students struggle with Maths and as a result avoid Math-related discipline or subjects despite the importance of Maths across many fields, including signal processing. Oftentimes, students' Maths difficulties stem from suboptimal pedagogy. We compared the MEGA method to the traditional step-by-step (CoT) method to ascertain which is better by using a within-group design after randomly assigning questions for the participants, who are university students. Samples (n=60) were randomly drawn from each of the two test sets of the Grade School Math 8K (GSM8K) and Mathematics Aptitude Test of Heuristics (MATH) datasets, based on the error margin of 11%, the confidence level of 90%, and a manageable number of samples for the student evaluators. These samples were used to evaluate two capable LLMs at length (Generative Pretrained Transformer 4o (GPT4o) and Claude 3.5 Sonnet) out of the initial six that were tested for capability. The results showed that students agree in more instances that the MEGA method is experienced as better for learning for both datasets. It is even much better than the CoT (47.5% compared to 26.67%) in the more difficult MATH dataset, indicating that MEGA is better at explaining difficult Maths problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。