arXiv:2505.10013cs.CL2025-05被引 2

提出DIF框架,量化大模型隐性偏见并发现准确率与偏见的反向关系。

DIF: A Framework for Benchmarking and Verifying Implicit Bias in LLMs

  • 用社会人口角色测试现有逻辑数学数据集,评估模型隐性偏见。
  • 发现模型答题准确率越高,隐性偏见越强,呈显著反向趋势。
  • 提供可解释的基准方法,适合关注模型公平性的研究者使用。

近年来,大型语言模型(LLMs)因其训练数据继承的潜在偏见而引发关注。以往研究揭示了模型在引入不同社会背景时响应变化的隐性偏见。我们认为,这种偏见不仅是伦理问题,更是技术问题,反映出模型无法有效处理外部信息。然而,与衡量模型智能的其他指标不同,目前尚无标准化方法来评估此类偏见。为此,我们提出一种可解释的基准方法——DIF(Demographic Implicit Fairness),通过在现有逻辑和数学问题数据集上引入社会人口角色进行评估,并结合零模型统计鲁棒性检验。实验表明该方法能有效验证隐性偏见的存在,并发现问答准确率与隐性偏见之间存在新颖的反向关系,支持我们的论点。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) have risen in prominence over the past few years, there has been concern over the potential biases in LLMs inherited from the training data. Previous studies have examined how LLMs exhibit implicit bias, such as when response generation changes when different social contexts are introduced. We argue that this implicit bias is not only an ethical, but also a technical issue, as it reveals an inability of LLMs to accommodate extraneous information. However, unlike other measures of LLM intelligence, there are no standard methods to benchmark this specific subset of LLM bias. To bridge this gap, we developed a method for calculating an easily interpretable benchmark, DIF (Demographic Implicit Fairness), by evaluating preexisting LLM logic and math problem datasets with sociodemographic personas, which is combined with a statistical robustness check using a null model. We demonstrate that this method can validate the presence of implicit bias in LLM behavior and find an novel inverse trend between question answering accuracy and implicit bias, supporting our argument.

大模型偏见公平性评估可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。