arXiv:2606.12422cs.CYcs.AI2026-06被引 1

用提示工程让大模型高效批改中小学生作业,数学科学评分接近真人。

Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering

论文配图:Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering
图 1 · 摘自论文原文
  • 用提示工程和上下文设计,让商用大模型按评分标准打分。
  • 数学与科学评分与真人一致度高,ELA表现差异明显。
  • 教师更信任AI的评语反馈,但对分数结果存疑,适合辅助教学。

大型语言模型(LLM)在教育评估中的应用正推动课堂评分方式的变革。尽管自动化评分系统已有数十年历史,生成式人工智能(GenAI)现可实现前所未有的效率与规模,支持基于标准的评分(SBG)。本文评估了一个利用商业基础模型并结合上下文与提示工程的LLM评分系统,针对数学、科学和英语语言艺术(ELA)科目进行评分。基于马萨诸塞州综合评估体系(MCAS)数据的实证研究显示,使用Claude Sonnet 4、Haiku 4.5、GPT-5和GPT-5 Mini模型,在数学和科学领域,其二次加权克朗巴赫系数(QWK)与均方误差降低比例(PRMSE)表现良好;而在ELA中表现各异。教师与学生反馈表明,对AI生成的评语接受度高,但对数值评分持保留态度。研究发现,融合AI效率与教师判断的混合模型能减轻工作负担、提升反馈质量,并支持公平评估,而不取代专业判断。

原文摘要 · Abstract (English)

The integration of large language models (LLMs) into educational assessment represents a transformative shift in classroom grading practices. While automated scoring systems and machine learning techniques have existed for decades, generative AI (GenAI) now enables educators to implement standards-based grading (SBG) with unprecedented efficiency and scale. This paper examines the theoretical foundations and evaluates an LLM grader that uses commercially available foundation models with context and prompt engineering to score student work against a rubric. Drawing on an empirical interrater agreement study using Massachusetts Comprehensive Assessment System (MCAS) data, we observed the Quadratic Weighted Kappa (QWK) and Proportional Reduction in Mean-Squared Error (PRMSE) across mathematics, science, and ELA, using Claude Sonnet 4, Haiku 4.5, GPT-5, and GPT-5 Mini. The results demonstrate that LLM graders, especially when based on foundational models with more parameters, achieve substantial agreement with human raters in mathematics and science assessments, while the performances vary in ELA, suggesting generic foundation models can be effective at scoring in given contexts. Additional analysis of teacher and student feedback reveals strong acceptance of AI-generated narrative feedback but skepticism toward numerical scores, suggesting that LLMs function most effectively as formative tools rather than summative evaluators. Our findings indicate that thoughtfully designed hybrid models that combine AI efficiency with teacher judgment can reduce workload, enhance feedback quality, and support equitable assessment practices without displacing professional expertise.

教育AI自动评分提示工程大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。