arXiv:2409.13843cs.CLcs.AI2024-09EMNLP被引 4

构建渐进式仇恨语数据集,评估大模型对偏见的敏感度。

STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions

  • 设计渐进式仇恨语数据集,覆盖9类人群46个子群体。
  • 顶尖模型检测偏见成功率仅19.3%至69.8%,表现不一。
  • 用人类判断对齐可提升敏感任务表现最高达191%。

缓解大型语言模型(LLMs)中的显性与隐性偏见已成为自然语言处理领域的关键挑战。然而,现有方法多孤立评估场景,未考虑上下文或偏见的连续性。为此,我们提出敏感性测试在仇恨升级(STOP)数据集,包含450个逐步升级的仇恨语例,共2,700个不同严重程度的句子,涵盖9类人群及46个子群体,确保全面性与包容性。我们评估了GPT-4、Mixtral和Llama 3等主流闭源与开源模型。结果显示,即使表现最佳的模型,其偏见检测一致性也差,准确率范围为19.3%至69.8%。此外,通过人类判断对齐模型,可在保持或提升性能的同时,使模型在BBQ、StereoSet和CrowS-Pairs等敏感任务上的回答率最高提升191%。STOP为评估大模型中复杂偏见提供了新框架,有助于制定更有效的偏见缓解策略,推动公平语言模型的发展。

原文摘要 · Abstract (English)

Mitigating explicit and implicit biases in Large Language Models (LLMs) has become a critical focus in the field of natural language processing. However, many current methodologies evaluate scenarios in isolation, without considering the broader context or the spectrum of potential biases within each situation. To address this, we introduce the Sensitivity Testing on Offensive Progressions (STOP) dataset, which includes 450 offensive progressions containing 2,700 unique sentences of varying severity that progressively escalate from less to more explicitly offensive. Covering a broad spectrum of 9 demographics and 46 sub-demographics, STOP ensures inclusivity and comprehensive coverage. We evaluate several leading closed- and open-source models, including GPT-4, Mixtral, and Llama 3. Our findings reveal that even the best-performing models detect bias inconsistently, with success rates ranging from 19.3% to 69.8%. We also demonstrate how aligning models with human judgments on STOP can improve model answer rates on sensitive tasks such as BBQ, StereoSet, and CrowS-Pairs by up to 191%, while maintaining or even improving performance. STOP presents a novel framework for assessing the complex nature of biases in LLMs, which will enable more effective bias mitigation strategies and facilitates the creation of fairer language models.

偏见检测大模型评估公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。