研究大模型对性别偏见的自信程度是否准确,发现多数模型存在偏差。
The Confidence Trap: Gender Bias and Predictive Certainty in LLMs
- 用新指标Gender-ECE衡量模型在性别代词消解任务中的自信度偏差
- 六款主流模型中,Gemma-2在性别偏见基准上校准最差
- 为伦理部署提供可信度评估方法,适合关注公平性的研究人员
大型语言模型(LLMs)在敏感领域应用增多,引发对其置信度评分与公平性、偏见关系的关注。本研究考察了模型预测置信度与人工标注的偏见判断之间的对齐情况,聚焦性别偏见,探究在涉及性别代词消解情境下的概率置信度校准问题。目标是评估基于预测置信度的校准指标能否有效捕捉模型中的公平性差异。结果显示,在六款前沿模型中,Gemma-2在性别偏见基准上的校准表现最差。本研究的主要贡献在于提出一种面向公平性的模型置信度校准评估方法,并引入新指标Gender-ECE,用于衡量性别代词解析任务中的性别差异。该工作为模型的伦理化部署提供了指导。
原文摘要 · Abstract (English)
The increased use of Large Language Models (LLMs) in sensitive domains leads to growing interest in how their confidence scores correspond to fairness and bias. This study examines the alignment between LLM-predicted confidence and human-annotated bias judgments. Focusing on gender bias, the research investigates probability confidence calibration in contexts involving gendered pronoun resolution. The goal is to evaluate if calibration metrics based on predicted confidence scores effectively capture fairness-related disparities in LLMs. The results show that, among the six state-of-the-art models, Gemma-2 demonstrates the worst calibration according to the gender bias benchmark. The primary contribution of this work is a fairness-aware evaluation of LLMs' confidence calibration, offering guidance for ethical deployment. In addition, we introduce a new calibration metric, Gender-ECE, designed to measure gender disparities in resolution tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。