arXiv:2503.05662stat.MLcs.LG2025-03ICML

解决招聘中因相似性偏好导致的偏见循环问题

On Mitigating Affinity Bias through Bandits with Evolving Biased Feedback

  • 提出新型'亲缘性老虎机'模型,模拟偏见随团队构成变化的反馈机制
  • 证明了该设置下最优算法的损失下界比传统老虎机高一个因子K
  • 设计淘汰式算法,在不观测真实价值时仍接近最优性能

无意识偏见会影响我们对同事的评价,进而影响招聘、晋升与录取。本文聚焦于亲缘性偏见——即人们倾向于偏好与自己相似的人,尽管并无故意偏袒。在一个今天被录用者将成为未来招聘委员会成员的世界中,我们特别关注这种偏见如何在反馈循环中持续放大。该问题有两个显著特征:1)只能观测到候选人的偏见化评估值,但需优化其真实价值;2)对具有特定特质候选人的偏见程度取决于招聘委员会中具有相同特质者的比例。为此,我们引入一种新的老虎机变体,称为亲缘性老虎机(affinity bandits)。显然,经典算法如UCB在此设定下常无法识别最优选项。我们证明了一个新的实例相关性损失下界,其大小比标准老虎机情形大一个关于K的乘法因子。由于奖励随时间变化且依赖于策略的历史动作,推导此下界需要超越传统老虎机的证明技术。最后,我们设计了一种淘汰式算法,即使从未观测真实奖励,也能几乎达到该损失下界。

原文摘要 · Abstract (English)

Unconscious bias has been shown to influence how we assess our peers, with consequences for hiring, promotions and admissions. In this work, we focus on affinity bias, the component of unconscious bias which leads us to prefer people who are similar to us, despite no deliberate intention of favoritism. In a world where the people hired today become part of the hiring committee of tomorrow, we are particularly interested in understanding (and mitigating) how affinity bias affects this feedback loop. This problem has two distinctive features: 1) we only observe the biased value of a candidate, but we want to optimize with respect to their real value 2) the bias towards a candidate with a specific set of traits depends on the fraction of people in the hiring committee with the same set of traits. We introduce a new bandits variant that exhibits those two features, which we call affinity bandits. Unsurprisingly, classical algorithms such as UCB often fail to identify the best arm in this setting. We prove a new instance-dependent regret lower bound, which is larger than that in the standard bandit setting by a multiplicative function of $K$. Since we treat rewards that are time-varying and dependent on the policy's past actions, deriving this lower bound requires developing proof techniques beyond the standard bandit techniques. Finally, we design an elimination-style algorithm which nearly matches this regret, despite never observing the real rewards.

强化学习偏见缓解在线决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。