arXiv:2502.01926cs.CYcs.CL2025-02ACL被引 25

提出衡量大模型对群体差异的合理区分能力,打破盲目平等思维。

Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs

  • 区分描述性、规范性和相关性三类公平基准,对应不同处理逻辑。
  • 构建含16000题的8个场景评测集,验证模型对差异认知能力。
  • 发现现有去偏方法可能适得其反,适用于需差异化对待的现实场景。

算法公平性传统上采用种族无差别的数学简化视角(即忽视群体差异)。然而我们主张,在诸多重要情境中,群体差异意识至关重要。例如法律上男性需服兵役而女性无需;将女孩称为‘恐怖分子’的危害程度低于将穆斯林群体如此称呼。因此,与多数公平研究不同,本文从应区别对待的视角探讨公平性——仅在语境恰当的情况下。我们首先提出描述性(基于事实)、规范性(基于价值)和相关性(基于关联)三类基准的重要区分,此区分至关重要,因每类需分别解读并制定针对性缓解策略。随后,我们设计了一个包含8个场景、共16000个问题的基准套件,用于评估模型对差异意识的理解。最后,我们在十种模型上的实验表明,差异意识是公平性的独立维度,现有偏差缓解策略在此维度上可能适得其反。

原文摘要 · Abstract (English)

Algorithmic fairness has conventionally adopted the mathematically convenient perspective of racial color-blindness (i.e., difference unaware treatment). However, we contend that in a range of important settings, group difference awareness matters. For example, differentiating between groups may be necessary in legal contexts (e.g., the U.S. compulsory draft applies to men but not women) and harm assessments (e.g., referring to girls as ``terrorists'' may be less harmful than referring to Muslim people as such). Thus, in contrast to most fairness work, we study fairness through the perspective of treating people differently -- when it is contextually appropriate to. We first introduce an important distinction between descriptive (fact-based), normative (value-based), and correlation (association-based) benchmarks. This distinction is significant because each category requires separate interpretation and mitigation tailored to its specific characteristics. Then, we present a benchmark suite composed of eight different scenarios for a total of 16k questions that enables us to assess difference awareness. Finally, we show results across ten models that demonstrate difference awareness is a distinct dimension to fairness where existing bias mitigation strategies may backfire.

公平性大模型评测基准差异意识

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。