用可解释的评分标准代替黑箱嵌入,减少标签偏差影响。
Mitigating Label Bias with Interpretable Rubric Embeddings
- 用专家定义的评分维度构建可解释嵌入表示
- 在硕士申请数据上降低群体差异并提升人才质量
- 适合需要公平性与透明度的招聘/招生场景
统计决策算法正被广泛应用于难以获取真实标签的领域,如招聘、大学录取和内容审核。此类系统常基于历史人工评估进行训练(例如以过往录用决定作为候选人能力的代理指标)。然而,若过去评估存在对特定群体的不公平倾向,模型可能继承这些偏见。为此,我们提出使用评分标准嵌入(rubric embeddings)来替代传统黑箱嵌入,其特征源自与目标构念一致的专家定义标准。通过将预测锚定在语义明确的维度上,该方法可防范有偏代理信号的影响。我们提供了理论与实证证据,证明在合理条件下,该方法能缓解标签偏差。我们在一个大型硕士项目申请的新数据集上进行了评估,结果表明,基于评分标准嵌入的模型在降低群体差异的同时提升了整体候选者质量。这表明,采用可解释且领域相关的表示方式,是应对有偏标签学习的有效路径。
原文摘要 · Abstract (English)
Statistical decision algorithms are increasingly deployed in domains where ground-truth labels are hard to obtain, such as hiring, university admissions, and content moderation. In these settings, models are typically trained on historical human evaluations -- for example, using past hiring decisions as a proxy for true applicant quality. However, if past evaluations unjustly favor certain groups, models trained on these labels may inherit those biases. To address this problem, we propose basing predictions on rubric embeddings, a representation framework that replaces standard black-box embeddings with features derived from expert-defined criteria that align with the underlying construct of interest. By anchoring predictions to semantically meaningful dimensions, this approach guards against biased proxy signals. We provide both theoretical and empirical evidence that rubric embeddings mitigate label bias under plausible conditions. Empirically, we evaluate our method on a novel dataset of applications to a large master's program. We find that models trained on rubric embeddings reduce group disparities while improving measures of cohort quality. Our results suggest that basing predictions on interpretable, domain-grounded representations offers a practical approach to learning in the presence of biased labels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。