改进偏见归因方法,首次用于菲律宾语等黏着语模型分析。
Bias Attribution in Filipino Language Models: Extending a Bias Interpretability Metric for Application on Agglutinative Languages
- 基于信息论改进偏见归因指标,适配黏着语语言结构。
- 发现菲律宾语模型偏见源于人、物、关系等实体主题。
- 为非英语语言模型的偏见分析提供新工具,适合语言学家和AI伦理研究者。
针对英文语言模型的偏见归因与可解释性研究已取得进展,我们在此基础上,将信息论偏见归因评分指标拓展至处理黏着语(如菲律宾语)的语言模型。通过在纯菲律宾语模型及三个多语言模型(一个全球训练、两个东南亚数据训练)上应用该方法,结果表明:菲律宾语模型的偏见主要由涉及人物、物体和关系的词汇驱动,这与英文模型中以行为(如犯罪、性、亲社会行为)为主的偏见主题形成鲜明对比。研究揭示了英语文本模型与非英语文本模型在社会人口群体相关输入处理上的根本差异。
原文摘要 · Abstract (English)
Emerging research on bias attribution and interpretability have revealed how tokens contribute to biased behavior in language models processing English texts. We build on this line of inquiry by adapting the information-theoretic bias attribution score metric for implementation on models handling agglutinative languages, particularly Filipino. We then demonstrate the effectiveness of our adapted method by using it on a purely Filipino model and on three multilingual models: one trained on languages worldwide and two on Southeast Asian data. Our results show that Filipino models are driven towards bias by words pertaining to people, objects, and relationships, entity-based themes that stand in contrast to the action-heavy nature of bias-contributing themes in English (i.e., criminal, sexual, and prosocial behaviors). These findings point to differences in how English and non-English models process inputs linked to sociodemographic groups and bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。