通过生成对抗性短语提升作文评分模型的公平性与鲁棒性
Phrase-Level Adversarial Training for Mitigating Bias in Neural Network-based Automatic Essay Scoring
- 在句子层面生成对抗样本,增强模型对偏见数据的抗性
- 实验表明新方法显著提升模型在对抗攻击下的表现
- 适合关注教育AI公平性的研究者与开发者
自动作文评分(AES)广泛用于教育评估,但因训练数据代表性不足,现有系统评分易受主流样本偏倚影响。本文提出一种无需依赖模型结构的短语级对抗训练方法,构建包含原始测试样本与对抗生成样本的攻击测试集。通过多种神经网络评分模型进行综合评估,结果表明该方法能显著提升模型在对抗样本存在及无攻击场景下的性能,有效缓解偏差并增强鲁棒性。
原文摘要 · Abstract (English)
Automatic Essay Scoring (AES) is widely used to evaluate candidates for educational purposes. However, due to the lack of representative data, most existing AES systems are not robust, and their scoring predictions are biased towards the most represented data samples. In this study, we propose a model-agnostic phrase-level method to generate an adversarial essay set to address the biases and robustness of AES models. Specifically, we construct an attack test set comprising samples from the original test set and adversarially generated samples using our proposed method. To evaluate the effectiveness of the attack strategy and data augmentation, we conducted a comprehensive analysis utilizing various neural network scoring models. Experimental results show that the proposed approach significantly improves AES model performance in the presence of adversarial examples and scenarios without such attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。