arXiv:2409.13705cs.CLcs.AI2024-09EMNLP被引 5

通过集成方法减轻文本安全分类器的偏见,提升公平性。

Debiasing Text Safety Classifiers through a Fairness-Aware Ensemble

  • 构建集成模型,统一多个分类器并减少偏见。
  • 在平衡数据上测试,公平性显著提升,性能基本不变。
  • 适合关注AI安全与公平性的研究人员和开发者。

大型语言模型(LLMs)广泛应用对输入输出的安全性保障提出更高要求。当安全分类器基于不平衡数据训练时,可能习得社会偏见。本文提出一种轻量级、后处理的去偏方法,通过构建集成模型,在不损害性能的前提下,不仅超越原始分类器表现,还能实现策略对齐并作为去偏正则化器。我们引入两种无需设定阈值的反事实公平性评估指标,并结合公平数据重加权(FDW)有效缓解偏见。我们扩充了Open AI数据集,并基于用户提示创建了一个新的模板化大模型生成数据集,两者均在身份群体间实现反事实平衡,覆盖四大安全领域;未来将公开发布。实验表明,该方法显著提升反事实公平性,对模型性能影响极小。

原文摘要 · Abstract (English)

Increasing use of large language models (LLMs) demand performant guardrails to ensure the safety of inputs and outputs of LLMs. When these safeguards are trained on imbalanced data, they can learn the societal biases. We present a light-weight, post-processing method for mitigating counterfactual fairness in closed-source text safety classifiers. Our approach involves building an ensemble that not only outperforms the input classifiers and policy-aligns them, but also acts as a debiasing regularizer. We introduce two threshold-agnostic metrics to assess the counterfactual fairness of a model, and demonstrate how combining these metrics with Fair Data Reweighting (FDW) helps mitigate biases. We create an expanded Open AI dataset, and a new templated LLM-generated dataset based on user-prompts, both of which are counterfactually balanced across identity groups and cover four key areas of safety; we will work towards publicly releasing these datasets. Our results show that our approach improves counterfactual fairness with minimal impact on model performance.

文本安全去偏公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。