为大模型解微积分题设计可解释框架,诊断其思维漏洞。
Interpretability Framework for LLMs in Undergraduate Calculus
- 提取推理流程并分解为带语义标签的步骤
- 发现模型解题虽语法流畅却常概念错误
- 适合教育AI研发者与课程设计者参考
大型语言模型在教育中应用日益广泛,但仅关注答案正确性无法衡量其解题质量、可靠性或教学有效性,尤其在数学领域,多步逻辑、符号推理和概念清晰度至关重要。传统评估方法主要关注最终答案准确率,忽视推理过程。为此,我们提出一种针对本科生微积分问题的新型可解释性框架。该方法结合推理流提取、将解题分解为语义标注的操作与概念,并通过提示词消融分析评估输入敏感性和输出稳定性。采用推理复杂度、短语敏感性和鲁棒性等结构化指标,在真实大学微积分一至三课程考试数据上评估模型行为。结果表明,大模型常生成语法流畅但概念错误的解答,其推理模式对提示语措辞和输入变化敏感。该框架支持细粒度的推理失败诊断,促进教学内容对齐,并指导可解释AI辅助反馈工具的设计。这是首个在数学教育中提供结构化、量化且具有教学基础的可解释性框架,为人工智能在STEM学习环境中的透明、负责任部署奠定基础。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly being used in education, yet their correctness alone does not capture the quality, reliability, or pedagogical validity of their problem-solving behavior, especially in mathematics, where multistep logic, symbolic reasoning, and conceptual clarity are critical. Conventional evaluation methods largely focus on final answer accuracy and overlook the reasoning process. To address this gap, we introduce a novel interpretability framework for analyzing LLM-generated solutions using undergraduate calculus problems as a representative domain. Our approach combines reasoning flow extraction and decomposing solutions into semantically labeled operations and concepts with prompt ablation analysis to assess input salience and output stability. Using structured metrics such as reasoning complexity, phrase sensitivity, and robustness, we evaluated the model behavior on real Calculus I to III university exams. Our findings revealed that LLMs often produce syntactically fluent yet conceptually flawed solutions, with reasoning patterns sensitive to prompt phrasing and input variation. This framework enables fine-grained diagnosis of reasoning failures, supports curriculum alignment, and informs the design of interpretable AI-assisted feedback tools. This is the first study to offer a structured, quantitative, and pedagogically grounded framework for interpreting LLM reasoning in mathematics education, laying the foundation for the transparent and responsible deployment of AI in STEM learning environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。