让大模型'代入'被偏见的群体,可有效缓解其固有偏见。
Persona Setting Pitfall: Persistent Outgroup Biases in Large Language Models Arising from Social Identity Adoption
- 通过引导模型代入被排斥群体视角,改变其认知立场。
- 模型对保守派和女性群体的偏见强度与对自由派的偏好相当。
- 为构建更公平的语言模型提供了可复制的新方法,适合伦理与AI安全研究者。
借鉴社会身份理论,我们探究了大语言模型(LLMs)如何内化由定向提示施加的身份。这些身份划分使模型产生“我们”(内群体)与“他们”(外群体)的区分,进而引发内群体偏爱和外群体偏见。然而,现有研究多聚焦于内群体偏爱,忽视了外群体偏见——这是群体间歧视的根本来源。本实验填补该空白,证明外群体偏见的强度与内群体偏爱相当。此外,通过引导模型采纳初始被贬低群体的视角,成功缓解了大模型中固有的亲自由派、反保守派偏见。该方法在性别偏见情境下亦得到复现。研究结果表明,可通过身份视角转换,发展更均衡、公正的语言模型。
原文摘要 · Abstract (English)
Drawing parallels between human cognition and artificial intelligence, we explored how large language models (LLMs) internalize identities imposed by targeted prompts. Informed by Social Identity Theory, these identity assignments lead LLMs to distinguish between "we" (the ingroup) and "they" (the outgroup). This self-categorization generates both ingroup favoritism and outgroup bias. Nonetheless, existing literature has predominantly focused on ingroup favoritism, often overlooking outgroup bias, which is a fundamental source of intergroup prejudice and discrimination. Our experiment addresses this gap by demonstrating that outgroup bias manifests as strongly as ingroup favoritism. Furthermore, we successfully mitigated the inherent pro-liberal, anti-conservative bias in LLMs by guiding them to adopt the perspectives of the initially disfavored group. These results were replicated in the context of gender bias. Our findings highlight the potential to develop more equitable and balanced language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。