arXiv:2601.17203cs.CL2026-01被引 24

用词向量偏差分析跨文化性别差距,揭示数据背后的现实不平等

Relating Word Embedding Gender Biases to Gender Gaps: A Cross-Cultural Analysis

  • 通过量化词向量中的性别偏见,反映真实社会性别差距
  • 在51个美国地区和99个国家的数据中验证了偏见与18项国际指标的相关性
  • 为理解文化背景下的性别不平等提供可量化的数据视角

现代自然语言处理模型常基于新闻、社交媒体等文化来源文本训练,近年被指出存在种族与性别偏见,这些偏见源于训练文本的固有倾向。尽管已有方法试图纠正此类偏见,但它们也可能反映生成文本的文化中实际存在的性别差距,从而帮助我们通过大数据理解文化背景。本文提出一种量化词向量性别偏见的方法,并利用其描述教育、政治、经济与健康领域的统计性别差距。我们在2018年覆盖51个美国地区和99个国家的推特数据上验证该方法,将州级与国家级的词向量偏见与18项国际及5项美国本土的性别差距统计数据进行相关性分析,揭示规律并评估预测能力。

原文摘要 · Abstract (English)

Modern models for common NLP tasks often employ machine learning techniques and train on journalistic, social media, or other culturally-derived text. These have recently been scrutinized for racial and gender biases, rooting from inherent bias in their training text. These biases are often sub-optimal and recent work poses methods to rectify them; however, these biases may shed light on actual racial or gender gaps in the culture(s) that produced the training text, thereby helping us understand cultural context through big data. This paper presents an approach for quantifying gender bias in word embeddings, and then using them to characterize statistical gender gaps in education, politics, economics, and health. We validate these metrics on 2018 Twitter data spanning 51 U.S. regions and 99 countries. We correlate state and country word embedding biases with 18 international and 5 U.S.-based statistical gender gaps, characterizing regularities and predictive strength.

性别偏见词向量跨文化分析大数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。