研究孟加拉语模型如何受数据、开发者和模型影响产生偏见
How do datasets, developers, and models affect biases in a low-resourced language?: The Case of the Bengali Language
- 对比mBERT与BanglaBERT在孟加拉语情感分析中的表现差异
- 发现不同身份类别下模型仍存在显著情感偏见
- 揭示跨开发者数据融合带来的认知不公与方法论挑战
社会技术系统如语言技术常表现出基于身份的偏见,加剧了历史上被边缘化群体的困境,且在低资源语境下研究不足。尽管针对特定语言或具备多语言支持的模型与数据集常被推荐以缓解偏见,本文在孟加拉语这一广泛使用但资源匮乏的语言中,实证检验了这些方法的有效性。我们对基于mBERT和BanglaBERT、在Google Dataset Search中可获取的所有孟加拉语情感分析(BSA)数据集上微调的情感分析模型进行了算法审计。分析显示,即使语义内容与结构相似,各类别模型仍存在跨性别、宗教与国籍的身份偏见。同时,我们考察了由不同人口背景个体创建的预训练模型与数据集组合所引发的不一致与不确定性,并将其与认识论不公、人工智能对齐及算法审计中的方法决策问题相联系。
原文摘要 · Abstract (English)
Sociotechnical systems, such as language technologies, frequently exhibit identity-based biases. These biases exacerbate the experiences of historically marginalized communities and remain understudied in low-resource contexts. While models and datasets specific to a language or with multilingual support are commonly recommended to address these biases, this paper empirically tests the effectiveness of such approaches in the context of gender, religion, and nationality-based identities in Bengali, a widely spoken but low-resourced language. We conducted an algorithmic audit of sentiment analysis models built on mBERT and BanglaBERT, which were fine-tuned using all Bengali sentiment analysis (BSA) datasets from Google Dataset Search. Our analyses showed that BSA models exhibit biases across different identity categories despite having similar semantic content and structure. We also examined the inconsistencies and uncertainties arising from combining pre-trained models and datasets created by individuals from diverse demographic backgrounds. We connected these findings to the broader discussions on epistemic injustice, AI alignment, and methodological decisions in algorithmic audits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。