AI助教可自动批改大学生数学自由作答,反馈质量接近人类专家。
Automated Feedback Generation for Undergraduate Mathematics: Development and Evaluation of an AI Teaching Assistant
- 基于大模型的模块化流程,支持自然语言输入与多维度评估。
- 在本科生作业上,生成反馈质量与人类专家相当。
- 可定制配置,适合数学课程教师快速部署使用。
智能辅导系统长期依赖结构化格式提供即时反馈,但对自由形式数学推理的可靠评估仍具挑战。本文提出一套系统,能处理自然语言输入,应对多种边界情况,并在技术正确性、写作风格与表达呈现方面给出专业评注。我们讨论了不同评估方法的优劣,结果显示,在所采用指标下,该系统生成的反馈质量与人类专家对早期本科生作业的评价相当。通过少量高阶和非常规题目进行压力测试,发现系统在复杂场景中存在明显短板,但也展现出令人鼓舞的成功案例。系统采用大语言模型构建模块化工作流,配置可读可编辑,无需编程知识,部分中间步骤可由教师预计算或注入。该工具已在帝国理工数学作业平台Lambdafeedback上线,并报告了集成情况。
原文摘要 · Abstract (English)
Intelligent tutoring systems have long enabled automated immediate feedback on student work when it is presented in a tightly structured format and when problems are very constrained, but reliably assessing free-form mathematical reasoning remains challenging. We present a system that processes free-form natural language input, handles a wide range of edge cases, and comments competently not only on the technical correctness of submitted proofs, but also on style and presentation issues. We discuss the advantages and disadvantages of various approaches to the evaluation of such a system, and show that by the metrics we evaluate, the quality of the feedback generated is comparable to that produced by human experts when assessing early undergraduate homework. We stress-test our system with a small set of more advanced and unusual questions, and report both significant gaps and encouraging successes in that more challenging setting. Our system uses large language models in a modular workflow. The workflow configuration is human-readable and editable without programming knowledge, and allows some intermediate steps to be precomputed or injected by the instructor. A version of our tool is deployed on the Imperial mathematics homework platform Lambdafeedback. We report also on the integration of our tool into this platform.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。