用课程标准指导LLM评分,让AI判卷更可信。
LLM-as-Judge in Education: A Curriculum-Grounded Marking Pipeline

- 基于课程大纲生成评分规则,确保判卷有据可依。
- AI评分结果与人工教师一致,且理由更可追溯。
- 适合教育机构开发标准化考试辅助系统。
生成式AI与大语言模型(LLMs)在题目生成和自动评估中应用日益广泛。然而,将LLM用于高风险考试准备不仅需要提示工程,还需构建系统化的软件流程,使模型输出严格基于教育主管部门发布的课程文件与评分指南。本文提出一种与工业伙伴联合开发的、基于课程的可配置LLM-as-Judge评分流水线,用于大学入学考试的题级评分。该流程识别问题涉及的主题、子主题及认知要求,并整合可验证、权威的上下文支持评分判断。课程意图通过具体课程资料实现,包括规定动词、学习目标、表现等级描述、术语定义及评分原则。采用分阶段的LLM工作流:先生成题目专属评分标准,明确表现预期;再推导并评估打分标准,以分配分数。该设计提升了评分一致性、透明度与官方评分实践的对齐性。初步评估显示,该流水线的评分结果与人类导师相当,且理由更可追溯至授权课程资料与评分标准。该系统已集成至在线学习平台,早期部署数据提供了运行使用与人工干预的初步洞察。
原文摘要 · Abstract (English)
Generative AI and large language models (LLMs) are increasingly applied to question generation and automated assessment. However, deploying LLMs in preparation for high-stakes exams requires more than prompt engineering; it demands software pipelines that systematically ground model outputs in authorised curriculum artefacts and marking guidelines issued by education authorities. This paper presents a curriculum-grounded, configurable LLM-as-Judge pipeline for question-level marking, co-developed with an industrial partner, to support exam preparation for university admission. The pipeline identifies the relevant topics, subtopics, and cognitive demand of a question, and assembles verifiable and authorised context to support LLM judgement. Curriculum intent is operationalised through concrete syllabus artefacts, including prescribed verbs and outcomes, performance band descriptors, glossary definitions, and marking-guideline principles. A staged LLM workflow is employed to first generate question-specific rubrics, capturing structured expectations of performance, and then derive and evaluate marking criteria used to allocate marks to student responses. This design improves consistency, transparency, and alignment with official marking practices. Preliminary evaluation shows that the proposed LLM-as-Judge pipeline delivers marking outcomes comparable to human tutors, while yielding justifications that are more traceable to authorised curriculum artefacts and marking standards. The pipeline has also been integrated into an online study platform, where early deployment data provide initial insights into operational usage and manual overrides.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。