arXiv:2601.03217cs.CL2026-01ACL被引 2

构建可执行的学生数学错题推理库,用于精准诊断学习者错误思维模式。

MalruleLib: Large-Scale Executable Misconception Reasoning with Step Traces for Modeling Student Thinking in Mathematics

  • 将67份教育研究中的误解转化为可执行程序,生成学生错误解题步骤
  • 9个大模型在跨模板预测中准确率从66%降至40%,验证认知偏差建模难度
  • 提供百万级生成数据与双路径推理追踪,适合教育AI与自适应学习系统

学生在数学中的错误往往是系统性的:学习者会一致地应用一种错误但连贯的解题步骤。我们提出MalruleLib,一个基于学习科学框架的工具,将67份学习科学和数学教育文献中的误解转化为可执行程序,并生成符合错误规则的逐步解题轨迹。我们将核心学生建模问题形式化为误算规则推理准确率(MRA):从一个错误解题案例推断出误解,并预测该学生在不同表述模板下的下一步答案。在九个语言模型(4B-120B参数)上,直接求解准确率为66%,而跨模板误解预测准确率下降至40%。MalruleLib编码了101条错误规则,覆盖498个可参数化的题目模板,生成正确推理与错误推理的双路径追踪。由于错误规则可执行、模板可扩展,系统可生成超百万实例,支持规模化监督与受控评估。实验发现跨模板性能下降10%-21%,而引入学生步骤轨迹可提升预测准确率3%-15%。我们发布MalruleLib作为教育人工智能基础设施,实现跨情境的学生过程建模,支持对根本误解的诊断与反馈。

原文摘要 · Abstract (English)

Student mistakes in mathematics are often systematic: a learner applies a coherent but wrong procedure and repeats it across contexts. We introduce MalruleLib, a learning-science-grounded framework that translates documented misconceptions into executable procedures, drawing on 67 learning-science and mathematics education sources, and generates step-by-step traces of malrule-consistent student work. We formalize a core student-modeling problem as Malrule Reasoning Accuracy (MRA): infer a misconception from one worked mistake and predict the student's next answer under cross-template rephrasing. Across nine language models (4B-120B), accuracy drops from 66% on direct problem solving to 40% on cross-template misconception prediction. MalruleLib encodes 101 malrules over 498 parameterized problem templates and produces paired dual-path traces for both correct reasoning and malrule-consistent student reasoning. Because malrules are executable and templates are parameterizable, MalruleLib can generate over one million instances, enabling scalable supervision and controlled evaluation. Using MalruleLib, we observe cross-template degradations of 10-21%, while providing student step traces improves prediction by 3-15%. We release MalruleLib as infrastructure for educational AI that models student procedures across contexts, enabling diagnosis and feedback that targets the underlying misconception.

教育AI错误建模认知诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。