研究大模型中姓名与职业的性别偏见相互影响机制。
On the Mutual Influence of Gender and Occupation in LLM Representations
- 分析姓名性别表征如何受职业性别刻板印象影响。
- 发现女性姓名嵌入更倾向关联女性主导职业,反之亦然。
- 揭示其在检测模型偏见中的潜力与局限性。
我们考察了大语言模型(LLMs)在不同职业语境下对姓名称谓的性别表征,探究职业与姓名性别感知在模型中的相互影响。研究发现,模型对姓名的性别表征与真实世界中该姓名关联的性别统计数据相关,并受到典型女性化或男性化职业共现的影响。此外,我们研究了姓名性别表征在下游职业预测任务中的影响,以及其作为内部指标识别模型外在偏见的潜力。结果显示,女性姓名嵌入通常提高女性主导职业的概率(男性姓名则相反),但将这些内部表征可靠用于偏见检测仍具挑战性。
原文摘要 · Abstract (English)
We examine LLM representations of gender for first names in various occupational contexts to study how occupations and the gender perception of first names in LLMs influence each other mutually. We find that LLMs' first-name gender representations correlate with real-world gender statistics associated with the name, and are influenced by the co-occurrence of stereotypically feminine or masculine occupations. Additionally, we study the influence of first-name gender representations on LLMs in a downstream occupation prediction task and their potential as an internal metric to identify extrinsic model biases. While feminine first-name embeddings often raise the probabilities for female-dominated jobs (and vice versa for male-dominated jobs), reliably using these internal gender representations for bias detection remains challenging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。