用人类与大模型混合群体降低偏见,效果优于纯模型群。
Wisdom from Diversity: Bias Mitigation Through Hybrid Human-LLM Crowds
- 混合人类与大模型的群体聚合,利用多样性缓解偏见。
- 局部加权聚合比平均法更有效,减少偏见并提升准确率。
- 人类提供多样性,大模型保证精度,适合敏感场景应用。
尽管性能优异,大语言模型(LLMs)仍可能延续训练数据中的偏见。通过对偏见诱导性标题的响应分析,发现模型常复现人类偏见。为应对此问题,我们探索基于群体的偏见缓解策略。结果表明,单纯平均多个大模型的输出会因群体多样性不足而加剧偏见;相比之下,局部加权聚合方法能更有效利用大模型群体的智慧,实现偏见降低与准确率提升。进一步发现,融合人类(多样性)与大模型(准确性)的混合群体,在种族与性别相关情境中显著提升性能并进一步减少偏见。
原文摘要 · Abstract (English)
Despite their performance, large language models (LLMs) can inadvertently perpetuate biases found in the data they are trained on. By analyzing LLM responses to bias-eliciting headlines, we find that these models often mirror human biases. To address this, we explore crowd-based strategies for mitigating bias through response aggregation. We first demonstrate that simply averaging responses from multiple LLMs, intended to leverage the "wisdom of the crowd", can exacerbate existing biases due to the limited diversity within LLM crowds. In contrast, we show that locally weighted aggregation methods more effectively leverage the wisdom of the LLM crowd, achieving both bias mitigation and improved accuracy. Finally, recognizing the complementary strengths of LLMs (accuracy) and humans (diversity), we demonstrate that hybrid crowds containing both significantly enhance performance and further reduce biases across ethnic and gender-related contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。