arXiv:2506.02425cs.CLstat.AP2025-06

用NLP分析22国英语教材,发现男性角色普遍更突出。

Gender Inequality in English Textbooks Around the World: an NLP Approach

  • 通过文本计数、首提顺序和词频关联量化性别不平等
  • 22国教材中男性角色在数量、首次提及和命名实体上均占优
  • 跨文化研究揭示全球普遍存在性别偏差,拉美区相对最均衡

教科书对儿童世界观塑造至关重要。尽管已有研究发现个别国家教材存在性别不平等,但跨文化比较仍较少。本研究采用自然语言处理方法,对来自7个文化圈的22个国家英语教材中的性别不平等现象进行量化分析。评估指标包括角色出现次数、首提顺序(firstness)以及基于TF-IDF的性别相关词汇关联度。研究还分析了TF-IDF词表中专有名词的性别分布模式,测试大语言模型区分性别化词汇列表的能力,并利用GloVe嵌入分析关键词与性别的关联紧密程度。结果显示,所有地区均存在性别不平等,男性角色在数量、首提顺序及命名实体方面显著占优;其中拉丁文化圈的性别差距最小。

原文摘要 · Abstract (English)

Textbooks play a critical role in shaping children's understanding of the world. While previous studies have identified gender inequality in individual countries' textbooks, few have examined the issue cross-culturally. This study applies natural language processing methods to quantify gender inequality in English textbooks from 22 countries across 7 cultural spheres. Metrics include character count, firstness (which gender is mentioned first), and TF-IDF word associations by gender. The analysis also identifies gender patterns in proper names appearing in TF-IDF word lists, tests whether large language models can distinguish between gendered word lists, and uses GloVe embeddings to examine how closely keywords associate with each gender. Results show consistent overrepresentation of male characters in terms of count, firstness, and named entities. All regions exhibit gender inequality, with the Latin cultural sphere showing the least disparity.

性别偏见NLP应用教育公平

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。