用共训练提升偏见检测,发现刻板印象是关键催化剂。
Stereotype Detection as a Catalyst for Enhanced Bias Detection: A Multi-Task Learning Approach
- 将偏见与刻板印象联合训练,提升检测效果。
- 联合训练使偏见检测准确率显著优于单独训练。
- 适合关注AI公平性与内容安全的研究者。
语言模型中的偏见与刻板印象可能在内容审核和决策等敏感领域造成伤害。本文探索联合学习偏见与刻板印象检测任务如何提升模型性能。我们提出 StereoBias 数据集,涵盖宗教、性别、社会经济地位、种族、职业等五个类别,支持对二者关系的深入研究。实验对比了编码器型模型与微调的解码器型模型(采用 QLoRA),结果显示编码器模型表现优异,解码器模型亦具竞争力。关键发现:联合训练偏见与刻板印象检测显著优于独立训练。额外实验表明,性能提升源于偏见与刻板印象间的内在关联,而非多任务学习本身。研究揭示利用刻板印象信息可构建更公平、高效的AI系统。
原文摘要 · Abstract (English)
Bias and stereotypes in language models can cause harm, especially in sensitive areas like content moderation and decision-making. This paper addresses bias and stereotype detection by exploring how jointly learning these tasks enhances model performance. We introduce StereoBias, a unique dataset labeled for bias and stereotype detection across five categories: religion, gender, socio-economic status, race, profession, and others, enabling a deeper study of their relationship. Our experiments compare encoder-only models and fine-tuned decoder-only models using QLoRA. While encoder-only models perform well, decoder-only models also show competitive results. Crucially, joint training on bias and stereotype detection significantly improves bias detection compared to training them separately. Additional experiments with sentiment analysis confirm that the improvements stem from the connection between bias and stereotypes, not multi-task learning alone. These findings highlight the value of leveraging stereotype information to build fairer and more effective AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。