用反事实实验发现教育大模型对性别反馈存在不对称偏差。
Benchmarking Educational LLMs with Analytics: A Case Study on Gender Bias in Feedback
- 通过替换文本中的性别词和提示中的性别背景,构造反事实测试样本。
- 六款模型中多数对男性描述产生更大语义偏移,仅部分响应显式性别提示。
- 揭示了模型在反馈中隐含的性别偏见,适合教育AI公平性研究者参考。
随着教师越来越多地在教学中使用生成式人工智能,我们需要可靠的基准方法来评估大语言模型(LLMs)在教学场景中的表现。本文提出一种基于嵌入的基准框架,用于检测生成性反馈中的偏见。利用来自AES 2.0语料库的600篇真实学生作文,我们构建了两种维度的控制反事实:(i) 通过词汇层面的性别词替换引入隐含线索;(ii) 通过提示中加入性别作者背景引入显式线索。研究涵盖六款代表性模型:GPT-5 mini、GPT-4o mini、DeepSeek-R1、DeepSeek-R1-Qwen、Gemini 2.5 Pro、Llama-3-8B。首先使用余弦距离与欧氏距离量化响应差异,再通过置换检验评估显著性,最后采用降维技术可视化结构。所有模型中,男性→女性反事实引发的语义变化大于女性→男性;仅有GPT与Llama模型对显式性别线索敏感。结果表明,即使最先进的模型在性别替换下仍表现出非对称的语义响应,暗示其反馈中存在持续的性别偏见。定性分析进一步显示,男性提示下反馈更具自主支持性,女性提示下则更显控制性。本文讨论了教育类GenAI公平审计的意义,提出学习分析中反事实评估的报告标准,并为提示设计与部署提供实践建议。
原文摘要 · Abstract (English)
As teachers increasingly turn to GenAI in their educational practice, we need robust methods to benchmark large language models (LLMs) for pedagogical purposes. This article presents an embedding-based benchmarking framework to detect bias in LLMs in the context of formative feedback. Using 600 authentic student essays from the AES 2.0 corpus, we constructed controlled counterfactuals along two dimensions: (i) implicit cues via lexicon-based swaps of gendered terms within essays, and (ii) explicit cues via gendered author background in the prompt. We investigated six representative LLMs (i.e. GPT-5 mini, GPT-4o mini, DeepSeek-R1, DeepSeek-R1-Qwen, Gemini 2.5 Pro, Llama-3-8B). We first quantified the response divergence with cosine and Euclidean distances over sentence embeddings, then assessed significance via permutation tests, and finally, visualised structure using dimensionality reduction. In all models, implicit manipulations reliably induced larger semantic shifts for male-female counterfactuals than for female-male. Only the GPT and Llama models showed sensitivity to explicit gender cues. These findings show that even state-of-the-art LLMs exhibit asymmetric semantic responses to gender substitutions, suggesting persistent gender biases in feedback they provide learners. Qualitative analyses further revealed consistent linguistic differences (e.g., more autonomy-supportive feedback under male cues vs. more controlling feedback under female cues). We discuss implications for fairness auditing of pedagogical GenAI, propose reporting standards for counterfactual evaluation in learning analytics, and outline practical guidance for prompt design and deployment to safeguard equitable feedback.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。