发现学术推荐信隐含性别线索,即使去标识后仍可被模型识别
Identifying and Mitigating Gender Cues in Academic Recommendation Letters: An Interpretability Case Study

- 用Transformer和LLM分析去性别化推荐信,发现性别泄露可达68%准确率
- 情感化、人道主义等词汇成女性推荐信的显著特征,是性别代理指标
- 移除这些线索后准确率仍高于随机,说明性别偏见根深蒂固
推荐信可能包含隐含的性别化语言,无意中影响招聘与录取决策。本文研究基于Transformer的编码器模型及大语言模型(LLMs)在去除姓名、代词等显式标识后,能否从美国医学住院医师项目提交的匿名推荐信中推断申请人性别。通过DistilBERT、RoBERTa和Llama 2三个模型对去性别化推荐信进行分类,最高达到68%的性别识别准确率。文本解释方法(如TF-IDF、SHAP)显示,“情感”“人道主义”等词汇常与女性申请者推荐信相关。通过移除这些隐含线索进行实验,重新训练分类器后准确率下降最多5.5%,宏F1分数下降2.7%,但预测性能仍显著优于随机水平。结果表明:1)推荐信中存在难以消除的性别标识线索,可能触发决策偏见;2)尽管技术框架提供公平评估的可行路径,仍需进一步探讨性别在推荐信评审中的作用。研究呼吁对真实场景中的评价文本进行上游审计,以补充模型层面的公平性干预。
原文摘要 · Abstract (English)
Letters of recommendation (LoRs) can carry patterns of implicitly gendered language that can inadvertently influence downstream decisions, e.g. in hiring and admissions. In this work, we investigate the extent to which Transformer-based encoder models as well as Large Language Models (LLMs) can infer the gender of applicants in academic LoRs submitted to an U.S. medical-residency program after explicit identifiers like names and pronouns are de-gendered. While using three models (DistilBERT, RoBERTa, and Llama 2) to classify the gender of anonymized and de-gendered LoRs, significant gender leakage was observed as evident from up to 68% classification accuracy. Text interpretation methods, like TF-IDF and SHAP, demonstrate that certain linguistic patterns are strong proxies for gender, e.g. "emotional'' and "humanitarian'' are commonly associated with LoRs from female applicants. As an experiment in creating truly gender-neutral LoRs, these implicit gender cues were remove resulting in a drop of up to 5.5% accuracy and 2.7% macro $F_1$ score on re-training the classifiers. However, applicant gender prediction still remains better than chance. In this case study, our findings highlight that 1) LoRs contain gender-identifying cues that are hard to remove and may activate bias in decision-making and 2) while our technical framework may be a concrete step toward fairer academic and professional evaluations, future work is needed to interrogate the role that gender plays in LoR review. Taken together, our findings motivate upstream auditing of evaluative text in real-world academic letters of recommendation as a necessary complement to model-level fairness interventions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。