arXiv:2606.18649cs.MAcs.CL2026-06

研究发现日本职场中大模型招聘存在亲女性偏见,名字是主要来源。

Gender Bias in LLM Hiring Decisions: Evidence from a Japanese Context and Evaluation of Mitigation Strategies

论文配图:Gender Bias in LLM Hiring Decisions: Evidence from a Japanese Context and Evaluation of Mitigation Strategies
图 1 · 摘自论文原文
  • 用60份日式履历和5个主流模型测试性别偏见
  • 所有模型均显示亲女性倾向,去除姓名可消除90%以上偏见
  • 提示词去性别化无效,隐私过滤器与GPT-4o冲突导致42%拒绝率

大型语言模型(LLMs)在招聘流程中的应用日益广泛,但现有研究多聚焦于英语语境下的西方简历。本研究探讨了亲女性偏见是否适用于日本企业环境,并评估两种实际缓解策略。基于60份日式履历(rirekisho格式),12组语言学上可信的性别信号姓名对,以及五种先进模型(Claude Sonnet 4.6、GPT-4o、DeepSeek-V3、Gemini 2.5 Flash、Llama 3.3 70B),在基线、提示指令和隐私过滤三种条件下共执行43,200次API调用。交叉随机效应线性混合模型确认所有五种模型均存在显著亲女性偏见,验证了非西方语境下该现象的普遍性。提示层面的去性别化指令未显著降低偏见。姓名依赖分析表明,候选人姓名是主要性别信号通道:从提示中移除姓名可使女性优势效应减少近90%。意外发现隐私过滤器与GPT-4o的内容安全机制不兼容,导致42%的请求被拒绝,凸显了姓名匿名化在实际部署中的挑战。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed in hiring workflows, yet most research on gender bias in LLM hiring decisions has focused on English-language, Western-format resumes. This study examines whether pro-female gender bias extends to a Japanese corporate context and evaluates two practical mitigation strategies. Using a counterfactual resume design with 60 Japanese rirekisho-format resumes, 12 name pairs selected on linguistically grounded gender-signal criteria, and five state-of-the-art LLMs (Claude Sonnet 4.6, GPT-4o, DeepSeek-V3, Gemini 2.5 Flash, Llama 3.3 70B), we conducted 43,200 API calls across baseline, prompt instruction, and privacy filter conditions. A crossed random-effects linear mixed model confirms a significant pro-female bias across all five models, replicating Western findings in a non-Western context. A prompt-level gender-neutrality instruction produces no meaningful reduction in bias. A name-reliance analysis formally identifies the candidate name as the primary gender channel: removing the name from the prompt reduces the female effect by nearly its full magnitude. An unexpected incompatibility between the privacy filter and GPT-4o's content safety filter, resulting in a 42% refusal rate, highlights a practical deployment challenge for name anonymization in LLM-assisted recruitment pipelines.

大模型偏见招聘算法性别公平日本语境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。