arXiv:2505.23996cs.CLcs.AI2025-05ICML被引 3

提出考虑模型不确定性的公平性评估方法,更精准发现大模型隐性偏见。

Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs

  • 引入新度量UCerF,结合模型置信度评估公平性。
  • 在31,756样本数据集上测试,发现Mistral-7B因高自信错误预测而不公平。
  • 适合关注模型透明性与可信AI的研究者使用。

大语言模型(LLMs)的快速普及凸显了对其公平性进行基准评测的重要性。传统公平性指标仅关注离散准确率(即预测正确性),未能捕捉模型不确定性带来的隐性影响(如对某群体的置信度高于另一群体,尽管准确率相似)。为解决此问题,我们提出一种考虑不确定性的公平性度量方法UCerF,实现更精细的公平性评估,能更好反映模型决策中的内在偏见。此外,鉴于现有数据集在数据量、多样性与清晰度上的不足,我们构建了一个包含31,756个样本的新性别-职业共指消解公平性评估数据集,更适合现代LLMs的评测。我们基于该度量与数据集建立基准,评估了十款开源大模型的表现。例如,Mistral-7B在错误预测上表现出过高置信度,导致公平性不佳,这一问题被等机会(Equalized Odds)忽略,但被UCerF识别。总体而言,本研究提出的具备不确定性感知的公平性基准,为发展更透明、可问责的AI系统铺平道路。

原文摘要 · Abstract (English)

The recent rapid adoption of large language models (LLMs) highlights the critical need for benchmarking their fairness. Conventional fairness metrics, which focus on discrete accuracy-based evaluations (i.e., prediction correctness), fail to capture the implicit impact of model uncertainty (e.g., higher model confidence about one group over another despite similar accuracy). To address this limitation, we propose an uncertainty-aware fairness metric, UCerF, to enable a fine-grained evaluation of model fairness that is more reflective of the internal bias in model decisions compared to conventional fairness measures. Furthermore, observing data size, diversity, and clarity issues in current datasets, we introduce a new gender-occupation fairness evaluation dataset with 31,756 samples for co-reference resolution, offering a more diverse and suitable dataset for evaluating modern LLMs. We establish a benchmark, using our metric and dataset, and apply it to evaluate the behavior of ten open-source LLMs. For example, Mistral-7B exhibits suboptimal fairness due to high confidence in incorrect predictions, a detail overlooked by Equalized Odds but captured by UCerF. Overall, our proposed LLM benchmark, which evaluates fairness with uncertainty awareness, paves the way for developing more transparent and accountable AI systems.

公平性评估大模型不确定性共指消解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。