arXiv:2602.23580cs.CLcs.AI2026-02中稿 · as a full paper at…

通过合成高分英语学习者样本,缓解自动化评分中的偏见放大问题。

BRIDGE the Gap: Mitigating Bias Amplification in Automated Scoring of English Language Learners via Inter-group Data Augmentation

  • 用非少数群体的高分内容填充英语学习者语言模式,生成高质量合成数据。
  • 在加州科学测试数据集上,显著降低高分英语学习者被低估的偏差,效果接近真实数据补充。
  • 适合教育公平、自动化评分与低资源场景下的算法改进研究者。

在教育评估领域,自动化评分系统越来越多地依赖深度学习和大语言模型。然而,这些系统存在偏见放大的风险:模型预测差距在学生群体间比训练数据中更大。这一问题对英语学习者(ELL)等少数群体尤为严重,因模型可能继承并加剧数据中的不平等。我们发现,这与表征偏差密切相关:少数群体(高分ELL)样本稀缺,使基于经验风险最小化的模型更偏向多数群体(非ELL)语言模式。因此,即使英语学习者具备相当的知识水平,模型仍会低估其得分,损害评分公平性。为此,我们提出BRIDGE——一种针对低资源评估场景的跨组数据生成框架。该方法不依赖有限的少数群体样本,而是将丰富的高分非ELL样本中与测评目标相关的知识内容,嵌入到真实的英语学习者语言模式中,生成高分样本。我们进一步引入判别器模型以保证合成样本质量。在加州科学测试(CAST)数据集上的实验表明,BRIDGE有效降低了高分英语学习者预测偏差,同时保持整体评分性能。值得注意的是,该方法实现的公平性提升可媲美额外真实人工数据,为大规模评估中的公平评分提供了一种低成本解决方案。

原文摘要 · Abstract (English)

In the field of educational assessment, automated scoring systems increasingly rely on deep learning and large language models (LLMs). However, these systems face significant risks of bias amplification, where model prediction gaps between student groups become larger than those observed in training data. This issue is especially severe for underrepresented groups such as English Language Learners (ELLs), as models may inherit and further magnify existing disparities in the data. We identify that this issue is closely tied to representation bias: the scarcity of minority (high-scoring ELL) samples makes models trained with empirical risk minimization favor majority (non-ELL) linguistic patterns. Consequently, models tend to under-predict ELL students who even demonstrate comparable domain knowledge but use different linguistic patterns, thereby undermining the fairness of automated scoring outcomes. To mitigate this, we propose BRIDGE, a Bias-Reducing Inter-group Data GEneration framework designed for low-resource assessment settings. Instead of relying on the limited minority samples, BRIDGE synthesizes high-scoring ELL samples by "pasting" construct-relevant (i.e., rubric-aligned knowledge and evidence) content from abundant high-scoring non-ELL samples into authentic ELL linguistic patterns. We further introduce a discriminator model to ensure the quality of synthetic samples. Experiments on California Science Test (CAST) datasets demonstrate that BRIDGE effectively reduces prediction bias for high-scoring ELL students while maintaining overall scoring performance. Notably, our method achieves fairness gains comparable to using additional real human data, offering a cost-effective solution for ensuring equitable scoring in large-scale assessments.

自动化评分公平性数据增强教育技术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。