arXiv:2410.02584cs.CLcs.CY2024-10EMNLP被引 61

研究多智能体大模型中的隐性性别偏见并提出有效缓解方法

Towards Implicit Bias Detection and Mitigation in Multi-Agent LLM Interactions

  • 构建场景数据集并设计评估指标检测隐性偏见
  • 多轮交互后偏见出现率超50%,且持续加剧
  • 结合自省与微调的组合策略效果最佳

随着大语言模型(LLMs)在社会模拟和各类社会任务中应用日益广泛,其因训练数据来自人类生成内容而易受社会偏见影响。本研究聚焦多智能体交互中隐性性别偏见的存在问题,提出两种缓解策略。首先构建可能引发隐性偏见的场景数据集,并开发评估指标。实证分析表明,LLMs输出中隐性偏见关联占比达50%以上,且在多智能体交互过程中进一步加剧。为此,提出基于上下文示例的自我反思(ICE)和监督微调两种方法。研究显示,二者均能有效降低偏见,尤其当两者结合时效果最优。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) continue to evolve, they are increasingly being employed in numerous studies to simulate societies and execute diverse social tasks. However, LLMs are susceptible to societal biases due to their exposure to human-generated data. Given that LLMs are being used to gain insights into various societal aspects, it is essential to mitigate these biases. To that end, our study investigates the presence of implicit gender biases in multi-agent LLM interactions and proposes two strategies to mitigate these biases. We begin by creating a dataset of scenarios where implicit gender biases might arise, and subsequently develop a metric to assess the presence of biases. Our empirical analysis reveals that LLMs generate outputs characterized by strong implicit bias associations (>= 50\% of the time). Furthermore, these biases tend to escalate following multi-agent interactions. To mitigate them, we propose two strategies: self-reflection with in-context examples (ICE); and supervised fine-tuning. Our research demonstrates that both methods effectively mitigate implicit biases, with the ensemble of fine-tuning and self-reflection proving to be the most successful.

多智能体偏见检测大模型伦理性别偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。