arXiv:2602.16438cs.LGcs.AI2026-02

研究目标性别对齐如何引发其他偏见的溢出效应,揭示大模型公平性评估的盲区。

Intra-Fairness Dynamics: The Bias Spillover Effect in Targeted LLM Alignment

  • 针对单一性别属性优化公平性,采用直接偏好优化与BBQ基准测试。
  • 在模糊语境下,外貌、性取向、残疾等属性偏见显著恶化(p<0.001)。
  • 强调需多属性、上下文感知的公平性评估,适用于大模型安全与伦理研究者。

传统大语言模型(LLM)公平性对齐主要关注单一敏感属性的偏见缓解,忽视了公平性本身具有多维性和情境依赖性。这种做法可能导致系统在特定维度上表现良好,却加剧了未被目标属性的偏见,即偏见溢出现象。尽管该现象在机器学习中已有研究,但在大模型对齐领域仍严重不足。本文研究了针对性别属性的对齐如何影响三个先进大模型(Mistral 7B、Llama 3.1 8B、Qwen 2.5 7B)在九个敏感属性上的公平性表现。通过使用直接偏好优化与BBQ基准,在模糊与明确语境下进行评估。结果发现:整体平均表现有所提升,但上下文感知分析显示,在模糊语境下,外貌(所有模型均p<0.001)、性取向和残疾状态等属性的偏见显著恶化。这表明,对某一属性的公平性优化可能在不确定性情境下无意中加剧其他属性的不平等,凸显了建立上下文感知、多属性公平性评估框架的必要性。

原文摘要 · Abstract (English)

Conventional large language model (LLM) fairness alignment largely focuses on mitigating bias along single sensitive attributes, overlooking fairness as an inherently multidimensional and context-specific value. This approach risks creating systems that achieve narrow fairness metrics while exacerbating disparities along untargeted attributes, a phenomenon known as bias spillover. While extensively studied in machine learning, bias spillover remains critically underexplored in LLM alignment. In this work, we investigate how targeted gender alignment affects fairness across nine sensitive attributes in three state-of-the-art LLMs (Mistral 7B, Llama 3.1 8B, Qwen 2.5 7B). Using Direct Preference Optimization and the BBQ benchmark, we evaluate fairness under ambiguous and disambiguous contexts. Our findings reveal noticeable bias spillover: while aggregate results show improvements, context-aware analysis exposes significant degradations in ambiguous contexts, particularly for physical appearance ($p< 0.001$ across all models), sexual orientation, and disability status. We demonstrate that improving fairness along one attribute can inadvertently worsen disparities in others under uncertainty, highlighting the necessity of context-aware, multi-attribute fairness evaluation frameworks.

大模型公平性偏见溢出多属性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。