研究对齐语言模型如何延续性别多元群体的伤害,发现主流方法会放大隐性偏见。
The Root Shapes the Fruit: On the Persistence of Gender-Exclusive Harms in Aligned Language Models
- 以跨性别者等群体为中心,评估对齐模型中的隐性偏见
- 16个DPO模型中均存在对性别多元群体的污名化与不认同语言
- 提出可推广的隐性偏见测量框架,适合其他社会群体
自然语言助手通过与人类偏好对齐来提供帮助性回应并避免有害输出。然而,对齐技术是否可能无意中延续甚至加剧其预对齐基础模型中的有害偏见,目前仍缺乏深入理解。现有偏见评估基准多聚焦于二元性别等主流社会类别,限制了对少数群体偏见的认知。为此,本文聚焦跨性别、非二元及其他性别多元身份,探究对齐过程如何与大模型中已存在的性别多元偏见相互作用。主要贡献包括:1)对领先偏好微调大模型的偏见评估范式进行系统调查,揭示性别多元代表性不足的关键缺口;2)对16个涵盖直接偏好优化(DPO)阶段的模型进行系统评估,发现主流基准未能检测到的性别多元群体真实伤害;3)提出一种适用于其他社会情境的隐性奖励信号有害偏见测量框架。结果表明,经DPO对齐的模型对监督微调(SFT)极为敏感,可能放大来自基模型的两种现实危害:污名化与非认同性语言。研究最后提出针对DPO及更广泛对齐实践的改进建议,倡导采用社区驱动的偏见评估框架,以更有效识别和应对少数群体的隐性伤害。
原文摘要 · Abstract (English)
Natural-language assistants are designed to provide users with helpful responses while avoiding harmful outputs, largely achieved through alignment to human preferences. Yet there is limited understanding of whether alignment techniques may inadvertently perpetuate or even amplify harmful biases inherited from their pre-aligned base models. This issue is compounded by the choice of bias evaluation benchmarks in popular preference-finetuned models, which predominantly focus on dominant social categories, such as binary gender, thereby limiting insights into biases affecting underrepresented groups. Towards addressing this gap, we center transgender, nonbinary, and other gender-diverse identities to investigate how alignment procedures interact with pre-existing gender-diverse bias in LLMs. Our key contributions include: 1) a comprehensive survey of bias evaluation modalities across leading preference-finetuned LLMs, highlighting critical gaps in gender-diverse representation, 2) systematic evaluation of gender-diverse biases across 16 models spanning Direct Preference Optimization (DPO) stages, uncovering harms popular bias benchmarks fail to detect, and 3) a flexible framework for measuring harmful biases in implicit reward signals applicable to other social contexts. Our findings reveal that DPO-aligned models are particularly sensitive to supervised finetuning (SFT), and can amplify two forms of real-world gender-diverse harms from their base models: stigmatization and gender non-affirmative language. We conclude with recommendations tailored to DPO and broader alignment practices, advocating for the adoption of community-informed bias evaluation frameworks to more effectively identify and address underrepresented harms in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。