揭示密集检索模型中性别偏差的内部机制,定位关键影响环节。
A Mechanistic Analysis of Gender Sensitivity in Dense Retrieval Models

- 发现性别敏感信号源于输入嵌入,经少数晚期注意力头传播。
- 嵌入层干预可泛化中和得分差异,注意力层干预能定向调整结果。
- 为精准去偏提供机制依据,适合关注模型公平性的研究者阅读。
尽管密集检索模型中的性别偏差已有广泛记录,以往研究显示模型常对男性相关文档评分高于女性或中性变体,但其内部产生差异的机制仍不清晰。本文对双编码器模型进行机制分析,定位到性别敏感性源自输入嵌入,并通过一组少量晚期注意力头传播,这些头同时携带性别与词项匹配信号。基于此发现,我们在两个识别出的关键点测试了引导干预,结果表明:嵌入层干预非特定地中和了得分差异,而注意力层干预则产生了方向性改变。研究为针对性去偏提供了机制基础,同时凸显在共享模型组件中分离性别与相关性信号的挑战。
原文摘要 · Abstract (English)
While gender bias in dense retrieval models is well documented, with prior work showing that models often score male-gendered documents higher than female or neutral variants, the internal mechanisms producing these disparities are poorly understood. In this paper, we mechanistically analyze bi-encoder models to localize gender sensitivity, finding that the signal originates in input embeddings and propagates through a small set of late-layer attention heads that carry both gender and term-matching signals. Guided by these findings, we test steering interventions at both identified points and find distinct effects: embedding-level steering non-specifically neutralizes score differences, while attention-level steering produces directional shifts. Our findings provide a mechanistic basis for targeted debiasing and highlight the challenge of disentangling gender from relevance signals in shared model components.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。