arXiv:2601.14506cs.CYcs.CL2026-01

检测到大模型在跨文化教育中对多重弱势群体的系统性偏见,差距达2.55个年级水平。

Compounding Disadvantage: Auditing Intersectional Bias in LLM-Generated Explanations Across Indian and American STEM Education

  • 用合成学生档案测试多模型在跨文化教育中的表现差异。
  • 弱势群体组合导致教育差距达2.55个年级,超出单一因素影响。
  • 偏见存在于所有模型且贯穿精英机构,需结构性审计。

大型语言模型正被广泛用于高收入与低收入国家的STEM教育中,提供个性化教学与反馈。这些系统旨在根据学生能力调整内容,但其是否基于真实能力或仅依赖人口特征尚未得到大规模验证。本研究发现,大模型生成的STEM内容在印度和美国两种文化背景下均系统性地偏向优势群体,最特权与最边缘化群体之间的差距高达2.55个年级水平。我们使用四种模型(Qwen 2.5-32B-Instruct、GPT-4o、GPT-4o-mini、GPT-OSS 20B),通过合成学生档案,交叉考察印度教育中的种姓、授课语言、高校层次,以及美国教育中的种族、非裔传统大学(HBCU)就读经历、学校类型,并结合收入、性别、残疾状况,在排序与生成任务中进行审计,采用FDR校正显著性检验与SHAP特征归因分析。结果显示:收入在所有模型和情境中均有显著影响;在印度语境中,授课语言产生最大单因素效应;残疾状态引发更简化的解释。偏见呈非加性复合:多重边缘化叠加带来的差距远超单一维度预测值,且偏见在顶尖机构中依然存在。该现象在四种架构中一致,且不随模型选择消除,表明在部署前必须开展跨文化、交叉性偏见审计。

原文摘要 · Abstract (English)

Large language models are increasingly deployed in STEM education for personalized instruction and feedback across institutions in high- and low-income countries. These systems are designed to adapt content to student needs, but whether they adapt based on demonstrated ability or demographic signals remains untested at scale. Here we establish that LLM-generated STEM content systematically disadvantages marginalized student profiles across two cultural contexts, with the gap between the most privileged and most marginalized profiles reaching 2.55 grade levels. We audited four LLMs (Qwen 2.5-32B-Instruct, GPT-4o, GPT-4o-mini, GPT-OSS 20B) using synthetic profiles crossing dimensions specific to Indian education (caste, medium of instruction, college tier) and American education (race, HBCU attendance, school type), alongside income, gender, and disability, across ranking and generation tasks with FDR-corrected significance testing and SHAP feature attribution. Income produces significant effects across every model and context, medium of instruction drives the largest single effect in the Indian context, and disability status triggers simpler explanations. Effects compound non-additively: marginalization across multiple dimensions produces gaps larger than any single dimension predicts, and biases persist within elite institutions. Bias is consistent across all four architectures and persists through model selection, making intersectional, cross-cultural auditing a structural requirement before deployment.

LLM偏见教育公平交叉性跨文化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。