用新指标发现大模型对中东存在系统性文化偏见。
Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS)
- 提出MECSS框架,将萨伊德的东方主义操作量化为可测维度。
- GPT-4和Falcon3-7B-Instruct均系统性再现东方主义结构,尤其是知识中心性偏移。
- 揭示‘萨伊德洗白’现象:模型否认概括却仍复制偏见,传统公平性评估无法捕捉。
AI系统正塑造数十亿人对异文化认知。当用户询问中东时,获得的并非中立事实,而是由训练数据中的西方话语框架所塑造的表征,而这些数据以英语为主。本文探究该表征是否符合萨伊德意义上的东方主义:否定了中东主体性、将西方框架视为中立而将非西方知识标记为特殊,并用外部范畴解释该地区。标准公平性度量无法识别此类结构性偏见,因它们只检测显性歧视。本文引入中东文化敏感性评分(MECSS),将萨伊德提出的七种东方主义操作转化为可测量维度,并定义“萨伊德洗白”——模型声称不进行泛化,却仍复现其所否认的结构。在280次对话(1,120次交互)中,GPT-4与Falcon3-7B-Instruct均系统性呈现东方主义模式,而非公开刻板印象。GPT-4平均得分为1.73,Falcon3-7B-Instruct为2.18,尽管后者在阿布扎比构建并使用阿拉伯语训练。这挑战了‘本地化建模即减少东方主义’的假设,但模型规模差异亦影响结果,地理因素难以孤立。认知中心性(Epistemic Center)在两模型中均接近最高分。萨伊德洗白出现在87.9%的GPT-4对话中,是现有度量无法察觉的模式。消除此类偏见需改变模型学习内容,而非仅增加语言或迁移机构。
原文摘要 · Abstract (English)
AI systems now shape how hundreds of millions of people learn about cultures other than their own. When someone asks one of these systems about the Middle East, they do not receive neutral facts. They receive a representation shaped by the frameworks embedded in training data, and that data is overwhelmingly Western and English-language. This paper asks whether that representation is Orientalist in Said's sense: whether it denies agency to Middle Eastern actors, treats Western frameworks as neutral while marking non-Western knowledge as particular, and explains the region through categories it did not produce. Standard fairness metrics cannot answer this, because they detect explicit prejudice rather than structural framing. This paper introduces the Middle East Cultural Sensitivity Score (MECSS), a framework that turns Said's seven Orientalist operations into measurable dimensions, and the term "Said-washing" for a specific failure: a model that disclaims generalization, then reproduces the structure it disclaimed. Across 280 conversations (1,120 exchanges), GPT-4 and Falcon3-7B-Instruct both reproduce Orientalist patterns systematically, through structural positioning rather than open stereotyping. GPT-4 scores moderately (mean MECSS 1.73); Falcon3-7B-Instruct scores higher (2.18), even though it was built in Abu Dhabi and trained with Arabic content. This is evidence against the assumption that building a model regionally makes it less Orientalist, though the models differ in size as well as origin, so geography cannot be isolated as the cause. Epistemic Center, the treatment of Western frameworks as unmarked universals, scores near the top of the scale for both models. Said-washing appears in 87.9% of GPT-4 conversations, a pattern existing metrics cannot see. Reducing this bias requires changing what models learn from, not only adding languages or relocating institutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。