GPT-4在跨文化规范推理中缺乏文化特异性,隐性偏见易被激活。
Is It Bad to Work All the Time? Cross-Cultural Evaluation of Social Norm Biases in GPT-4
- 通过叙事推理评估GPT-4对不同文化的理解能力
- 模型生成的规范普遍缺乏文化差异性,隐性刻板印象可被轻松恢复
- 揭示了大模型在跨文化公平性上的深层偏差,适合关注AI伦理的研究者
大型语言模型(LLMs)已被证明与西方或北美文化价值观保持一致。以往研究主要依赖直接提问(过去是人类,现在也包括模型)来衡量价值观。然而,很难相信这些模型在真实场景中能持续遵循这些价值观。为此,我们采用自下而上的方法,要求模型基于来自不同文化的叙事进行规范推理。结果发现,GPT-4倾向于生成虽不必然错误但显著缺乏文化特异性的规范。尽管避免显性生成刻板印象,某些文化隐性刻板印象并未被真正抑制,仍可轻易被恢复。解决这些问题对于开发能公平服务多元用户群体的LLM至关重要。
原文摘要 · Abstract (English)
LLMs have been demonstrated to align with the values of Western or North American cultures. Prior work predominantly showed this effect through leveraging surveys that directly ask (originally people and now also LLMs) about their values. However, it is hard to believe that LLMs would consistently apply those values in real-world scenarios. To address that, we take a bottom-up approach, asking LLMs to reason about cultural norms in narratives from different cultures. We find that GPT-4 tends to generate norms that, while not necessarily incorrect, are significantly less culture-specific. In addition, while it avoids overtly generating stereotypes, the stereotypical representations of certain cultures are merely hidden rather than suppressed in the model, and such stereotypes can be easily recovered. Addressing these challenges is a crucial step towards developing LLMs that fairly serve their diverse user base.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。