探究英语模型去偏方法在马耳他语模型中的迁移效果
From Measurement to Mitigation: Exploring the Transferability of Debiasing Approaches to Gender Bias in Maltese Language Models
- 将英语去偏技术适配到马耳他语的单/多语言BERT模型
- 发现现有方法在形态复杂的低资源语言中效果有限
- 为马耳他语性别偏见研究提供新评估数据集
大型语言模型(LLMs)虽显著推动自然语言处理发展,但其仍易受社会偏见影响,尤其在训练数据中反映有害刻板印象,对边缘群体造成不公。本文首次测量马耳他语语言模型中的性别偏见,强调此类偏见会强化社会成见并忽视性别多样性,尤其在语法性别复杂、资源匮乏的语言中更为严重。针对英语主流模型已有的偏见评估与去偏方法,本文将其适配至马耳他语的 BERTu(单语)和 mBERTu(多语)模型,采用 CrowS-Pairs 与 SEAT 等基准测试,并应用反事实数据增强、丢弃正则化、Auto-Debias 和 GuiDebias 等去偏技术。研究还贡献了新的马耳他语性别偏见评估数据集。结果表明,现有去偏方法在语言结构复杂的低资源语言中迁移困难,凸显了构建更具包容性的多语言NLP体系的必要性。
原文摘要 · Abstract (English)
The advancement of Large Language Models (LLMs) has transformed Natural Language Processing (NLP), enabling performance across diverse tasks with little task-specific training. However, LLMs remain susceptible to social biases, particularly reflecting harmful stereotypes from training data, which can disproportionately affect marginalised communities. We measure gender bias in Maltese LMs, arguing that such bias is harmful as it reinforces societal stereotypes and fails to account for gender diversity, which is especially problematic in gendered, low-resource languages. While bias evaluation and mitigation efforts have progressed for English-centric models, research on low-resourced and morphologically rich languages remains limited. This research investigates the transferability of debiasing methods to Maltese language models, focusing on BERTu and mBERTu, BERT-based monolingual and multilingual models respectively. Bias measurement and mitigation techniques from English are adapted to Maltese, using benchmarks such as CrowS-Pairs and SEAT, alongside debiasing methods Counterfactual Data Augmentation, Dropout Regularization, Auto-Debias, and GuiDebias. We also contribute to future work in the study of gender bias in Maltese by creating evaluation datasets. Our findings highlight the challenges of applying existing bias mitigation methods to linguistically complex languages, underscoring the need for more inclusive approaches in the development of multilingual NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。