arXiv:2606.08400cs.SEcs.AI2026-06中稿 · ISET 2026被引 1

研究大模型在研究生课程评分中的表现,发现其评分标准会随使用历史偏移。

Impacts of Histories and Models on LLM Grading: A Study in Advanced Software Engineering Courses

  • 设计人机协同评分流程,用大模型辅助批改180份研究生报告
  • 不同模型间评分差异大,且连续使用导致评分标准逐渐偏离人工标准
  • 简单合并模型结果无法提升与人工评分的一致性,需规范操作

研究生阶段的研究阅读报告评估给教师带来巨大工作负担。尽管大语言模型(LLMs)在自动化学术评分方面潜力巨大,但其在这一专业任务中的可靠性仍缺乏研究,尤其是评分一致性不足,已成为教育公平的主要障碍。本文提出一种人机对齐的LLM辅助评分流程,并基于某高级软件工程课程的180份学生提交作业开展案例研究。我们评估了两种主流LLM——Grok和GPT——在评分一致性和与人工评分对齐度方面的表现。结果表明,模型内部存在不同程度的一致性,而不同模型间评分差异显著;简单的集成方法无法提升与人工评价的一致性。关键发现是:持续交互历史会引发模型评分标准系统性偏离人类专家评分。研究证明了大模型在减轻研究生教育评阅负担方面的潜力,同时指出盲目使用可能导致系统性不公平,建议需制定特定操作规范以缓解此类偏差。

原文摘要 · Abstract (English)

Graduate-level research reading report assessment creates a substantial labor burden for educators. While large language models (LLMs) hold great potential for automating academic grading, their reliability for this specialized task remains understudied, particularly regarding grading consistency, the lack of which represents a primary obstacle to educational fairness. This paper proposes a human-aligned LLM-assisted grading workflow and presents a case study based on 180 student submissions from a graduate advanced software engineering course. We evaluate two mainstream LLMs, Grok and GPT, in terms of grading consistency and alignment with human scores. We find LLMs exhibit distinct levels of intra-model consistency and significant inter-model grading inconsistencies, while simple ensemble approaches cannot improve alignment with human evaluation. Critically, continuous interaction history drives systematic drift in models' grading standards away from human expert scores. Our findings demonstrate LLMs' potential in reducing grading workload for educators in graduate education, while highlighting that indiscriminate LLM grading may introduce systemic unfairness, suggesting that specific operational practices are required to mitigate such disparities.

大模型评分教育公平一致性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。