用新数据集和多种方法,让大模型更公平地减少九类社会偏见。
Do the Right Thing, Just Debias! Multi-Category Bias Mitigation Using LLMs
- 构建1507对语句的ANUBIS数据集,覆盖九类社会偏见。
- T5模型通过微调、强化学习等方法有效降低多类别偏见。
- 结果可跨数据集迁移,助力开发更公平的AI系统。
本文针对语言模型中鲁棒且通用的偏见缓解难题,提出新型数据集ANUBIS,包含1507对精心筛选的句子对,涵盖九类社会偏见。评估了T5等前沿模型在监督微调(SFT)、强化学习(PPO、DPO)及上下文学习(ICL)下的偏见缓解效果。研究聚焦于多类别社会偏见削减、跨数据集泛化能力以及训练模型的环境影响。ANUBIS数据集与研究成果为构建更公平的AI系统提供了宝贵资源,推动负责任、无偏见技术的发展,具有广泛的社会意义。
原文摘要 · Abstract (English)
This paper tackles the challenge of building robust and generalizable bias mitigation models for language. Recognizing the limitations of existing datasets, we introduce ANUBIS, a novel dataset with 1507 carefully curated sentence pairs encompassing nine social bias categories. We evaluate state-of-the-art models like T5, utilizing Supervised Fine-Tuning (SFT), Reinforcement Learning (PPO, DPO), and In-Context Learning (ICL) for effective bias mitigation. Our analysis focuses on multi-class social bias reduction, cross-dataset generalizability, and environmental impact of the trained models. ANUBIS and our findings offer valuable resources for building more equitable AI systems and contribute to the development of responsible and unbiased technologies with broad societal impact.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。