GPT-4在印度种姓与宗教叙事中存在深层刻板印象,即使提示多样化也难纠正。
How Deep Is Representational Bias in LLMs? The Cases of Caste and Religion
- 通过7200+故事生成测试,对比真实人口分布评估模型偏见
- 主导群体被严重高估,偏差程度远超训练数据分布
- 反复提示调整效果有限,表明需根本性模型改进
大型语言模型(LLMs)中的表征偏见主要通过单次响应交互测量,且集中于欧美背景的身份维度如种族与性别。本文系统审计GPT-4 Turbo,探究其在印度种姓与宗教等较少研究身份维度上的偏见深度。我们设计多种提示,要求生成超过7,200个关于重要人生事件(如婚礼)的故事,以激发多样性。将输出中宗教与种姓的分布与印度人口普查数据对比,量化模型在这些维度上的表征偏见及其“顽固性”。结果发现,尽管提示旨在促进多样性,但模型仍持续高估文化主导群体,远超其实际人口比例。此外,偏见呈现“赢家通吃”特征,比训练数据可能存在的分布偏见更严重,且多次提示干预效果有限且不一致。这表明,仅靠多样化训练数据无法有效矫正模型偏见,亟需更根本性的模型开发策略。数据集与代码本:https://github.com/agrimaseth/How-Deep-Is-Representational-Bias-in-LLMs
原文摘要 · Abstract (English)
Representational bias in large language models (LLMs) has predominantly been measured through single-response interactions and has focused on Global North-centric identities like race and gender. We expand on that research by conducting a systematic audit of GPT-4 Turbo to reveal how deeply encoded representational biases are and how they extend to less-explored dimensions of identity. We prompt GPT-4 Turbo to generate over 7,200 stories about significant life events (such as weddings) in India, using prompts designed to encourage diversity to varying extents. Comparing the diversity of religious and caste representation in the outputs against the actual population distribution in India as recorded in census data, we quantify the presence and "stickiness" of representational bias in the LLM for religion and caste. We find that GPT-4 responses consistently overrepresent culturally dominant groups far beyond their statistical representation, despite prompts intended to encourage representational diversity. Our findings also suggest that representational bias in LLMs has a winner-take-all quality that is more biased than the likely distribution bias in their training data, and repeated prompt-based nudges have limited and inconsistent efficacy in dislodging these biases. These results suggest that diversifying training data alone may not be sufficient to correct LLM bias, highlighting the need for more fundamental changes in model development. Dataset and Codebook: https://github.com/agrimaseth/How-Deep-Is-Representational-Bias-in-LLMs
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。