arXiv:2603.18765cs.CL2026-03被引 1

LLM自动评分受文风影响,写作不正式会被大幅扣分。

Implicit Grading Bias in Large Language Models: How Writing Style Affects Automated Assessment Across Math, Programming, and Essay Tasks

  • 用控制变量法测试不同文风对评分影响
  • 非正式表达平均扣1.9分,接近成绩降档
  • 数学编程题无明显偏见,作文题最敏感

随着大型语言模型在教育场景中被用于自动评分,其评估的公平性与偏差问题日益突出。本研究探究当内容正确性保持不变时,语言模型是否因写作风格产生隐性评分偏差。我们构建了涵盖数学、编程和作文三类任务的180份学生作答数据集,每类任务包含语法错误、非正式语言、非母语表达三种表面扰动类型。使用两款先进开源模型LLaMA 3.3 70B(Meta)和Qwen 2.5 72B(Alibaba),在明确指令要求仅评估内容正确性、忽略文风的前提下进行评分(1-10分制)。结果显示,在作文任务中,两个模型对所有扰动类型均存在显著评分偏差(p < 0.05),效应量从中等(Cohen's d = 0.64)到极大(d = 4.25)。非正式语言被重罚,LLaMA平均扣1.90分,Qwen扣1.20分,相当于从B+降至C+;非母语表达分别扣1.35和0.90分。而数学与编程任务则几乎无偏差,多数条件未达统计显著性。结果表明,语言模型评分偏差具有任务依赖性、风格敏感性,且即便在提示中明确反偏,偏差仍持续存在。研究呼吁在机构采用前建立偏差审计机制。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly deployed as automated graders in educational settings, concerns about fairness and bias in their evaluations have become critical. This study investigates whether LLMs exhibit implicit grading bias based on writing style when the underlying content correctness remains constant. We constructed a controlled dataset of 180 student responses across three subjects (Mathematics, Programming, and Essay/Writing), each with three surface-level perturbation types: grammar errors, informal language, and non-native phrasing. Two state-of-the-art open-source LLMs -- LLaMA 3.3 70B (Meta) and Qwen 2.5 72B (Alibaba) -- were prompted to grade responses on a 1-10 scale with explicit instructions to evaluate content correctness only and to disregard writing style. Our results reveal statistically significant grading bias in Essay/Writing tasks across both models and all perturbation types (p < 0.05), with effect sizes ranging from medium (Cohen's d = 0.64) to very large (d = 4.25). Informal language received the heaviest penalty, with LLaMA deducting an average of 1.90 points and Qwen deducting 1.20 points on a 10-point scale -- penalties comparable to the difference between a B+ and C+ letter grade. Non-native phrasing was penalized 1.35 and 0.90 points respectively. In sharp contrast, Mathematics and Programming tasks showed minimal bias, with most conditions failing to reach statistical significance. These findings demonstrate that LLM grading bias is subject-dependent, style-sensitive, and persists despite explicit counter-bias instructions in the grading prompt. We discuss implications for equitable deployment of LLM-based grading systems and recommend bias auditing protocols before institutional adoption.

自动评分语言模型公平性教育技术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。