用大模型评估学生编程解题过程,精准判断代数能力
LLM-Driven Rubric-Based Assessment of Algebraic Competence in Multi-Stage Block Coding Tasks with Design and Field Evaluation
- 用大模型按五维评分标准分析编码过程
- 42名初中生测试中与专家评分高度一致
- 适合在线数学教育平台做过程性评价
随着在线教育平台发展,亟需既能衡量答案正确性,又能反映学生认知深度的评估方法。本研究提出并验证了一种基于大语言模型(LLM)的评分框架,用于评估真实情境下的块状编程任务中的代数能力。题目由数学教育专家设计,每个环节对应五个预设评分维度,使LLM能同时评估结果正确性与解题过程质量。系统部署于在线平台,记录所有中间响应,并由LLM进行符合评分标准的学业表现评估。为检验实际效果,开展了包含42名初中生的实地研究,参与多阶段二次方程求解任务。研究结合学习者自评与专家评分,作为基准对比系统输出。结果显示,基于LLM的评分与专家判断高度一致,且持续生成符合评分标准的过程反馈。这证明了该框架在在线数学与STEM教育平台中的有效性与可扩展性。
原文摘要 · Abstract (English)
As online education platforms continue to expand, there is a growing need for assessment methods that not only measure answer accuracy but also capture the depth of students' cognitive processes in alignment with curriculum objectives. This study proposes and evaluates a rubric-based assessment framework powered by a large language model (LLM) for measuring algebraic competence, real-world-context block coding tasks. The problem set, designed by mathematics education experts, aligns each problem segment with five predefined rubric dimensions, enabling the LLM to assess both correctness and quality of students' problem-solving processes. The system was implemented on an online platform that records all intermediate responses and employs the LLM for rubric-aligned achievement evaluation. To examine the practical effectiveness of the proposed framework, we conducted a field study involving 42 middle school students engaged in multi-stage quadratic equation tasks with block coding. The study integrated learner self-assessments and expert ratings to benchmark the system's outputs. The LLM-based rubric evaluation showed strong agreement with expert judgments and consistently produced rubric-aligned, process-oriented feedback. These results demonstrate both the validity and scalability of incorporating LLM-driven rubric assessment into online mathematics and STEM education platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。