arXiv:2506.02740cs.CL2025-06被引 28

从网络文本中提取性别刻板印象行为,验证其与人类认知的匹配度。

Stereotypical gender actions can be extracted from Web text

  • 利用常识库和用户性别信息,从文本中挖掘性别化行为
  • 预测结果与人类判断相关性达0.47,分类准确率AUC为0.76
  • 提供441个带人工标注的性别倾向动作数据集

我们从文本语料库和推特数据中提取了具有性别特征的行为,并与人们的刻板期待进行对比。采用开放心智常识(OMCS)知识库聚焦于与常识和日常生活相关的动作。通过推特用户性别信息及基于代词/姓名的性别启发式方法,计算动作的性别偏见。在高召回率下,语料库预测与人类黄金标准的斯皮尔曼相关系数达到0.47,预测黄金标准极性的受试者工作特征曲线下面积(AUC)为0.76。研究证实,自然语言(尤其是推特语料)可用于补充常识知识库中的行为性别刻板印象。此外,本文还发布了两个数据集:441个常识动作的人工评分数据集,以及21,442个由本研究方法自动评分的动作数据集。

原文摘要 · Abstract (English)

We extracted gender-specific actions from text corpora and Twitter, and compared them to stereotypical expectations of people. We used Open Mind Common Sense (OMCS), a commonsense knowledge repository, to focus on actions that are pertinent to common sense and daily life of humans. We use the gender information of Twitter users and Web-corpus-based pronoun/name gender heuristics to compute the gender bias of the actions. With high recall, we obtained a Spearman correlation of 0.47 between corpus-based predictions and a human gold standard, and an area under the ROC curve of 0.76 when predicting the polarity of the gold standard. We conclude that it is feasible to use natural text (and a Twitter-derived corpus in particular) in order to augment commonsense repositories with the stereotypical gender expectations of actions. We also present a dataset of 441 commonsense actions with human judges' ratings on whether the action is typically/slightly masculine/feminine (or neutral), and another larger dataset of 21,442 actions automatically rated by the methods we investigate in this study.

性别偏见常识推理文本挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。