提出新指标解释语言模型偏见来源,发现东南亚多语模型存在性别与性向偏见
A Novel Interpretability Metric for Explaining Bias in Language Models: Applications on Multilingual Models from Southeast Asia
- 基于信息论设计词级偏见归因分数,量化每个词对偏见的贡献
- 发现东南亚多语模型存在显著性别与性向偏见,尤其在犯罪、亲密关系等话题中
- 适用于关注模型公平性与偏见解释的研究者,尤其适合多语种场景
预训练语言模型(PLMs)的偏见研究多聚焦于偏见评估与缓解,却忽视了偏见归因与可解释性问题。本文提出一种新指标——偏见归因分数,基于信息论衡量语言模型中词级对偏见行为的贡献。通过应用于包括东南亚多语模型在内的多个模型,验证了该指标的有效性。结果表明,这些模型中存在显著的性别与性向偏见。可解释性与语义分析进一步揭示,模型偏见主要由涉及犯罪、亲密关系、助人等话语类别中的词汇引发,说明这些领域是模型从预训练数据中复制偏见的高风险区,使用时需格外谨慎。
原文摘要 · Abstract (English)
Work on bias in pretrained language models (PLMs) focuses on bias evaluation and mitigation and fails to tackle the question of bias attribution and explainability. We propose a novel metric, the $\textit{bias attribution score}$, which draws from information theory to measure token-level contributions to biased behavior in PLMs. We then demonstrate the utility of this metric by applying it on multilingual PLMs, including models from Southeast Asia which have not yet been thoroughly examined in bias evaluation literature. Our results confirm the presence of sexist and homophobic bias in Southeast Asian PLMs. Interpretability and semantic analyses also reveal that PLM bias is strongly induced by words relating to crime, intimate relationships, and helping among other discursive categories, suggesting that these are topics where PLMs strongly reproduce bias from pretraining data and where PLMs should be used with more caution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。