arXiv:2409.13884cs.CLcs.AI2024-09被引 15

用多个大模型协作,有效降低语言模型的偏见。

A Multi-LLM Debiasing Framework

  • 设计中心化与去中心化两种多模型协作机制
  • 在多个社会群体上显著减少模型偏见
  • 适合关注AI公平性与可解释性的研究者

大型语言模型(LLMs)虽有巨大社会潜力,却普遍存在加剧社会不平等的偏见。尽管已有数据增强、零样本提示和微调等缓解方法,但偏见仍持续存在,包括人类难以察觉的细微偏见。近期研究显示,多大模型协作能提升推理质量与事实准确性。本文提出首个系统评估两种路径的多大模型去偏框架:中心化模式由单一核心模型主导对话,去中心化模式则各模型直接交互。实验表明,该框架显著降低模型偏见,在多个社会群体上优于基线方法。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are powerful tools with the potential to benefit society immensely, yet, they have demonstrated biases that perpetuate societal inequalities. Despite significant advancements in bias mitigation techniques using data augmentation, zero-shot prompting, and model fine-tuning, biases continuously persist, including subtle biases that may elude human detection. Recent research has shown a growing interest in multi-LLM approaches, which have been demonstrated to be effective in improving the quality of reasoning and factuality in LLMs. Building on this approach, we propose a novel multi-LLM debiasing framework aimed at reducing bias in LLMs. Our work is the first to introduce and evaluate two distinct approaches within this framework for debiasing LLMs: a centralized method, where the conversation is facilitated by a single central LLM, and a decentralized method, where all models communicate directly. Our findings reveal that our multi-LLM framework significantly reduces bias in LLMs, outperforming the baseline method across several social groups.

去偏见多模型协作LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。