GPT模型输出受训练数据意识形态影响,语言差异暴露其政治倾向性。
Identifying the sources of ideological bias in GPT models through linguistic variation in output
- 通过对比不同语言的输出风格,识别训练数据中的意识形态偏见。
- 波兰语输出更保守,瑞典语更自由,且此差异在GPT-3.5和GPT-4中均存在。
- 过滤策略无法消除根本偏见,高质量数据集才是关键。
现有研究显示,GPT-3.5和GPT-4等生成式AI模型会延续社会刻板印象与偏见。一个较少被关注的偏见来源是意识形态。这些模型是否会在政治敏感话题上持特定立场?本文提出一种新方法,通过分析使用不同政治态度国家的语言,评估GPT在敏感政治话题上的平均输出。结果发现:在与保守社会匹配的语言(如波兰语)中,模型输出更保守;在仅用于自由社会的语言(如瑞典语)中,输出更自由。该现象有力证明了训练数据中的偏见。此外,尽管GPT-4因OpenAI的过滤政策更偏向自由,但语言间的差异仍持续存在。核心结论是:生成模型训练应聚焦高质量、经过筛选的数据集,即使需牺牲部分数据规模。训练后的过滤仅引入新偏见,无法消除底层训练偏见。
原文摘要 · Abstract (English)
Extant work shows that generative AI models such as GPT-3.5 and 4 perpetuate social stereotypes and biases. One concerning but less explored source of bias is ideology. Do GPT models take ideological stances on politically sensitive topics? In this article, we provide an original approach to identifying ideological bias in generative models, showing that bias can stem from both the training data and the filtering algorithm. We leverage linguistic variation in countries with contrasting political attitudes to evaluate bias in average GPT responses to sensitive political topics in those languages. First, we find that GPT output is more conservative in languages that map well onto conservative societies (i.e., Polish), and more liberal in languages used uniquely in liberal societies (i.e., Swedish). This result provides strong evidence of training data bias in GPT models. Second, differences across languages observed in GPT-3.5 persist in GPT-4, even though GPT-4 is significantly more liberal due to OpenAI's filtering policy. Our main takeaway is that generative model training must focus on high-quality, curated datasets to reduce bias, even if it entails a compromise in training data size. Filtering responses after training only introduces new biases and does not remove the underlying training biases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。