arXiv:2411.07656cs.CLcs.MA2024-11被引 1

用协作智能体改善大模型对跨性别者代词使用偏见

Mitigating Bias in Queer Representation within Large Language Models: A Collaborative Agent Approach

  • 设计多智能体系统,分步检测并修正代词使用偏见
  • 在Tango数据集上使包容性代词识别率提升32.6个百分点
  • 适合关注AI公平性与社会负责任生成内容的研究者

大型语言模型常在代词使用中延续偏见,导致跨性别群体被误指或排除。本文聚焦于模型输出中不当使用传统性别代词(如'他''她')的问题,提出一种协作智能体框架,通过专门的检测与修正智能体实现代词使用的优化。基于面向性别代词使用的基准数据集Tango的实验表明,该方法在正确拒绝不恰当传统代词方面,相比GPT-4o提升了32.6个百分点(χ² = 38.57, p < 0.0001)。结果凸显了智能体驱动框架在提升AI生成内容公平性与包容性方面的潜力,证明其在减少偏见、推动社会责任型AI方面的有效性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) often perpetuate biases in pronoun usage, leading to misrepresentation or exclusion of queer individuals. This paper addresses the specific problem of biased pronoun usage in LLM outputs, particularly the inappropriate use of traditionally gendered pronouns ("he," "she") when inclusive language is needed to accurately represent all identities. We introduce a collaborative agent pipeline designed to mitigate these biases by analyzing and optimizing pronoun usage for inclusivity. Our multi-agent framework includes specialized agents for both bias detection and correction. Experimental evaluations using the Tango dataset-a benchmark focused on gender pronoun usage-demonstrate that our approach significantly improves inclusive pronoun classification, achieving a 32.6 percentage point increase over GPT-4o in correctly disagreeing with inappropriate traditionally gendered pronouns $(χ^2 = 38.57, p < 0.0001)$. These results accentuate the potential of agent-driven frameworks in enhancing fairness and inclusivity in AI-generated content, demonstrating their efficacy in reducing biases and promoting socially responsible AI.

大模型偏见代词公平智能体协作社会影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。