arXiv:2411.10636cs.CLcs.AI2024-11

针对孟加拉语模型的外在性别偏见,提出有效缓解方法并公开数据集。

Mitigating Extrinsic Gender Bias for Bangla Classification Tasks

  • 通过替换性别词汇构造带偏见的数据对,实现精准评估偏见程度。
  • 新方法RandSymKL在降低偏见的同时保持分类准确率领先。
  • 首个面向孟加拉语分类任务的系统性偏见评估与缓解框架,适合低资源语言研究者。

本研究探讨了孟加拉语预训练语言模型中的外在性别偏见,这一领域在低资源语言中尚未被充分探索。为评估偏见,我们构建了四个手工标注的任务特定基准数据集,涵盖情感分析、毒性检测、仇恨言论检测和反讽检测。每个数据集均通过细微的性别扰动进行增强,系统性地替换性别相关姓名和称谓,同时保持语义不变,从而实现最小差异对评估性别驱动的预测变化。为此,我们提出一种名为RandSymKL的随机化去偏策略,结合对称KL散度与交叉熵损失,统一整合到任务特定预训练模型中,以缓解外在性别偏见。该方法在多个基准上优于现有去偏技术,在显著降低偏见的同时保持了竞争力的准确率。为促进后续研究,我们已公开代码与数据集:https://github.com/sajib-kumar/Mitigating-Bangla-Extrinsic-Gender-Bias。

原文摘要 · Abstract (English)

In this study, we investigate extrinsic gender bias in Bangla pretrained language models, a largely underexplored area in low-resource languages. To assess this bias, we construct four manually annotated, task-specific benchmark datasets for sentiment analysis, toxicity detection, hate speech detection, and sarcasm detection. Each dataset is augmented using nuanced gender perturbations, where we systematically swap gendered names and terms while preserving semantic content, enabling minimal-pair evaluation of gender-driven prediction shifts. We then propose RandSymKL, a randomized debiasing strategy integrated with symmetric KL divergence and cross-entropy loss to mitigate the bias across task-specific pretrained models. RandSymKL is a refined training approach to integrate these elements in a unified way for extrinsic gender bias mitigation focused on classification tasks. Our approach was evaluated against existing bias mitigation methods, with results showing that our technique not only effectively reduces bias but also maintains competitive accuracy compared to other baseline approaches. To promote further research, we have made both our implementation and datasets publicly available: https://github.com/sajib-kumar/Mitigating-Bangla-Extrinsic-Gender-Bias

性别偏见孟加拉语去偏自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。