探究大模型文化偏见:越贴近全球多样性,越易输出歧视性内容
Should LLMs be WEIRD? Exploring WEIRDness and Human Rights in Large Language Models
- 用世界价值观调查数据测试五款主流大模型的文化倾向
- 非西方模型响应更多元,但人权违规率高出2%-4%
- 适合关注伦理风险与跨文化AI的从业者阅读
大型语言模型(LLMs)通常基于反映WEIRD价值观(西方、受教育、工业化、富裕、民主)的数据训练,引发文化偏见与公平性担忧。本文利用世界价值观调查数据,评估了GPT-3.5、GPT-4、Llama-3、BLOOM和Qwen五款主流模型,分析其回应与WEIRD国家价值观的契合度,以及是否违背人权原则。为体现全球多样性,结果与《世界人权宣言》及亚洲、中东、非洲三个区域宪章对比。结果显示,与WEIRD价值观契合度较低的模型(如BLOOM、Qwen)虽生成更多文化多样性回应,但其输出违反人权的比例高出2%至4%,尤其在性别与平等议题上,部分模型认同如“无法生育的男人不是真正男人”、“丈夫应知道妻子行踪”等有害性别观念。这表明,提升文化代表性可能加剧歧视性内容生成风险。尽管宪法式AI等方法可嵌入人权准则,但难以完全化解此矛盾。
原文摘要 · Abstract (English)
Large language models (LLMs) are often trained on data that reflect WEIRD values: Western, Educated, Industrialized, Rich, and Democratic. This raises concerns about cultural bias and fairness. Using responses to the World Values Survey, we evaluated five widely used LLMs: GPT-3.5, GPT-4, Llama-3, BLOOM, and Qwen. We measured how closely these responses aligned with the values of the WEIRD countries and whether they conflicted with human rights principles. To reflect global diversity, we compared the results with the Universal Declaration of Human Rights and three regional charters from Asia, the Middle East, and Africa. Models with lower alignment to WEIRD values, such as BLOOM and Qwen, produced more culturally varied responses but were 2% to 4% more likely to generate outputs that violated human rights, especially regarding gender and equality. For example, some models agreed with the statements ``a man who cannot father children is not a real man'' and ``a husband should always know where his wife is'', reflecting harmful gender norms. These findings suggest that as cultural representation in LLMs increases, so does the risk of reproducing discriminatory beliefs. Approaches such as Constitutional AI, which could embed human rights principles into model behavior, may only partly help resolve this tension.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。