用大模型生成符合教学标准的个性化数学应用题。
EDUMATH: Generating Standards-aligned Educational Math Word Problems

- 结合专家与大模型评估,构建首个教师标注的标准对齐数学题数据集。
- 120亿参数模型性能媲美更大模型,300亿参数模型零训练超越闭源基线。
- 学生测试显示生成题与人工题效果相当,但更受学生欢迎。
数学应用题(MWPs)是中小学教育的关键工具,根据学生兴趣和能力定制可提升学习效果。但教师因班级规模大、工作压力高,难以实现个性化定制。我们提出利用大语言模型(LLMs)生成符合教育标准且个性化的数学应用题。通过联合人类专家与LLM评判,评估了超过11,000道由开源与闭源模型生成的题目,并构建了首个教师标注的标准对齐教育数学题数据集。我们利用该数据训练一个120亿参数的开源模型,其性能达到更大更强开源模型水平。此外,使用该数据训练文本分类器,使300亿参数开源模型在无需微调的情况下超越现有闭源基线。我们还发现,本模型生成的题目比现有模型更接近人工编写风格。最后,首次开展面向小学学生的定制化生成题实验,结果显示学生在生成题与人工题上的表现相当,但更偏好生成题。
原文摘要 · Abstract (English)
Math word problems (MWPs) are critical K-12 educational tools, and customizing them to students' interests and ability levels can enhance learning. However, teachers struggle to find time to customize MWPs for students given large class sizes and increasing burnout. We propose that LLMs can support math education by generating MWPs customized to student interests and math education standards. We use a joint human expert-LLM judge approach to evaluate over 11,000 MWPs generated by open and closed LLMs and develop the first teacher-annotated dataset for standards-aligned educational MWP generation. We show the value of our data by using it to train a 12B open model that matches the performance of larger and more capable open models. We also use our teacher-annotated data to train a text classifier that enables a 30B open LLM to outperform existing closed baselines without any training. Next, we show our models' MWPs are more similar to human-written MWPs than those from existing models. We conclude by conducting the first study of customized LLM-generated MWPs with grade school students, finding they perform similarly on our models' MWPs relative to human-written MWPs but consistently prefer our customized MWPs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。