研究大模型在交叉身份下的公平性,发现其表现受刻板印象影响严重。
Intersectional Fairness in Large Language Models
- 通过模糊与明确语境测试六款大模型的公平性表现
- 模型在符合刻板印象时准确率更高,尤其在种族性别交叉群体中
- 现有评估需结合偏差、子群公平性和重复实验一致性
大型语言模型(LLMs)在社会敏感场景中广泛应用,引发对其公平性和偏见的担忧,尤其是在交叉性人口特征方面。本文使用两个基准数据集中的模糊和明确语境,系统评估了六款LLM的交叉公平性。通过偏差分数、子群公平性指标、准确率及多轮重复实验中正负极性问题的一致性分析模型行为。结果显示,尽管现代LLM在模糊语境中表现良好,但因非未知预测稀疏,公平性度量信息有限;在明确语境中,模型准确率受刻板印象对齐影响,当正确答案强化刻板印象时更准确,尤其在种族-性别交叉情境下方向性偏见更明显。子群公平性指标进一步显示,尽管某些情况下观测到差异低,但结果分布仍不均衡。多次运行中响应也存在不一致,包括刻板印象一致的回答。总体而言,模型看似的能力部分源于刻板印象一致线索,且无一模型在交叉情境中表现出持续可靠或公平的行为。研究强调需超越准确率评估,主张结合偏差、子群公平性与重复实验一致性,在交叉群体、语境及多次运行中综合衡量。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly deployed in socially sensitive settings, raising concerns about fairness and biases, particularly across intersectional demographic attributes. In this paper, we systematically evaluate intersectional fairness in six LLMs using ambiguous and disambiguated contexts from two benchmark datasets. We assess LLM behavior using bias scores, subgroup fairness metrics, accuracy, and consistency through multi-run analysis across contexts and negative and non-negative question polarities. Our results show that while modern LLMs generally perform well in ambiguous contexts, this limits the informativeness of fairness metrics due to sparse non-unknown predictions. In disambiguated contexts, LLM accuracy is influenced by stereotype alignment, with models being more accurate when the correct answer reinforces a stereotype than when it contradicts it. This pattern is especially pronounced in race-gender intersections, where directional bias toward stereotypes is stronger. Subgroup fairness metrics further indicate that, despite low observed disparity in some cases, outcome distributions remain uneven across intersectional groups. Across repeated runs, responses also vary in consistency, including stereotype-aligned responses. Overall, our findings show that apparent model competence is partly associated with stereotype-consistent cues, and no evaluated LLM achieves consistently reliable or fair behavior across intersectional settings. These findings highlight the need for evaluation beyond accuracy, emphasizing the importance of combining bias, subgroup fairness, and consistency metrics across intersectional groups, contexts, and repeated runs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。