arXiv:2505.02456cs.CL2025-05被引 8

检测大模型在跨性别跨国家职业推荐中的交织偏见

Colombian Waitresses y Jueces canadienses: Gender and Country Biases in Occupation Recommendations from LLMs

  • 构建英西德三语提示基准,系统测试25国与4类性别组合
  • 多模型测试显示:即使单维度公平,交叉偏见仍显著存在
  • 指令微调模型偏差最低且最稳定,提示语语言影响显著

公平性研究常聚焦英语和单一性别偏见,本文首次系统考察大语言模型在多语言环境下国家与性别交织的就业推荐偏见。构建包含英、西、德三语的基准数据集,涵盖25个国家和4组代词,评估5个基于Llama的模型。结果表明:尽管个别维度表现平衡,但国家与性别交叉导致的职业推荐偏差依然严重;提示语言显著影响偏见水平,指令微调模型表现出最低且最稳定的偏见水平。研究呼吁公平性评估需采用多语言与交叉视角。

原文摘要 · Abstract (English)

One of the goals of fairness research in NLP is to measure and mitigate stereotypical biases that are propagated by NLP systems. However, such work tends to focus on single axes of bias (most often gender) and the English language. Addressing these limitations, we contribute the first study of multilingual intersecting country and gender biases, with a focus on occupation recommendations generated by large language models. We construct a benchmark of prompts in English, Spanish and German, where we systematically vary country and gender, using 25 countries and four pronoun sets. Then, we evaluate a suite of 5 Llama-based models on this benchmark, finding that LLMs encode significant gender and country biases. Notably, we find that even when models show parity for gender or country individually, intersectional occupational biases based on both country and gender persist. We also show that the prompting language significantly affects bias, and instruction-tuned models consistently demonstrate the lowest and most stable levels of bias. Our findings highlight the need for fairness researchers to use intersectional and multilingual lenses in their work.

大模型偏见多语言交叉性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。