用大模型自动优化评分标准,让机器评分更接近真人。
A LLM-Powered Automatic Grading Framework with Human-Level Guidelines Optimization
- 构建多智能体系统,用大模型自我反思纠错来优化评分规则。
- 在教师知识题上准确率超越主流方法,评分行为更贴近人类。
- 适合需要高一致性评分的教育评测场景,尤其对开放题适用。
开放型简答题(SAGs)在学习分析中被广泛用于深入理解学习者作答内容,但其评估面临工作量大和评分不一致的问题。随着自然语言处理的发展,自动简答题评分(ASAG)成为潜在解决方案。然而,现有方法泛化能力差,常针对特定题目定制。本文提出统一的多智能体框架GradeOpt,利用大语言模型作为评分者,并引入反思者与优化者两个额外大模型代理,实现对原始评分标准的自动优化。通过在教学内容知识(PCK)与学科内容知识(CK)题目的挑战性任务上的实验,GradeOpt在评分准确率和与人类评分行为的一致性上均优于代表性基线。全面的消融实验证明了各组件的有效性。
原文摘要 · Abstract (English)
Open-ended short-answer questions (SAGs) have been widely recognized as a powerful tool for providing deeper insights into learners' responses in the context of learning analytics (LA). However, SAGs often present challenges in practice due to the high grading workload and concerns about inconsistent assessments. With recent advancements in natural language processing (NLP), automatic short-answer grading (ASAG) offers a promising solution to these challenges. Despite this, current ASAG algorithms are often limited in generalizability and tend to be tailored to specific questions. In this paper, we propose a unified multi-agent ASAG framework, GradeOpt, which leverages large language models (LLMs) as graders for SAGs. More importantly, GradeOpt incorporates two additional LLM-based agents - the reflector and the refiner - into the multi-agent system. This enables GradeOpt to automatically optimize the original grading guidelines by performing self-reflection on its errors. Through experiments on a challenging ASAG task, namely the grading of pedagogical content knowledge (PCK) and content knowledge (CK) questions, GradeOpt demonstrates superior performance in grading accuracy and behavior alignment with human graders compared to representative baselines. Finally, comprehensive ablation studies confirm the effectiveness of the individual components designed in GradeOpt.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。