大模型作文评分系统对无关因素有较强鲁棒性,但抄袭会降分。
Measuring What Matters -- or What's Convenient?: Robustness of LLM-Based Scoring Systems to Construct-Irrelevant Factors
- 用双架构大模型评估情境判断题短文,关注无关特征影响。
- 填充无意义文本或拼写错误不影响得分,抄袭反而降低分数。
- 适合关注AI评分公平性与设计优化的研究者参考。
自动化评分系统在教育测评领域广泛应用,尤其在开放式作答和作文评分中表现接近或优于人工评分员。然而,这些系统常受无关因素(即与测评目标无关的响应特征)及对抗性条件影响。随着大语言模型在自动化评分中的普及,其幻觉问题和对无关因素的鲁棒性受到关注。本研究考察了一种双架构大语言模型评分系统在情境判断测试中对短文作答的评分表现,发现该系统对添加无意义文本、拼写错误和写作复杂度具有普遍鲁棒性;而重复大段文字则导致平均得分下降,与非大模型系统的结论相反;偏离主题的回答被严重扣分。结果表明,只要注重构念相关性,未来大模型评分系统具备良好鲁棒性。
原文摘要 · Abstract (English)
Automated systems have been widely adopted across the educational testing industry for open-response assessment and essay scoring. These systems commonly achieve performance levels comparable to or superior than trained human raters, but have frequently been demonstrated to be vulnerable to the influence of construct-irrelevant factors (i.e., features of responses that are unrelated to the construct assessed) and adversarial conditions. Given the rising usage of large language models in automated scoring systems, there is a renewed focus on ``hallucinations'' and the robustness of these LLM-based automated scoring approaches to construct-irrelevant factors. This study investigates the effects of construct-irrelevant factors on a dual-architecture LLM-based scoring system designed to score short essay-like open-response items in a situational judgment test. It was found that the scoring system was generally robust to padding responses with meaningless text, spelling errors, and writing sophistication. Duplicating large passages of text resulted in lower scores predicted by the system, on average, contradicting results from previous studies of non-LLM-based scoring systems, while off-topic responses were heavily penalized by the scoring system. These results provide encouraging support for the robustness of future LLM-based scoring systems when designed with construct relevance in mind.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。