arXiv:2411.17338cs.CLcs.AI2024-11中稿 · NeurIPS被引 5

用真实统计数据评估大模型偏见,发现不同标准下偏见表现不同。

Different Bias Under Different Criteria: Assessing Bias in LLMs with a Fact-Based Approach

  • 基于真实人口数据构建客观偏见评估指标
  • 实验证明人类更倾向符合现实分布的模型输出
  • 强调需多视角评估,避免单一标准误导

大型语言模型常反映现实世界中的偏见,现有方法多以群体均等对待为无偏标准,但因对平等的理解差异难以形成统一标准。本文提出一种基于事实的评估方法,利用真实世界统计数据作为客观基准。通过人类调查发现,当模型输出与现实人口分布一致时,人类感知更积极。在多种大模型上应用该指标显示,模型偏见程度随评估标准变化,凸显多维度评估的必要性。

原文摘要 · Abstract (English)

Large language models (LLMs) often reflect real-world biases, leading to efforts to mitigate these effects and make the models unbiased. Achieving this goal requires defining clear criteria for an unbiased state, with any deviation from these criteria considered biased. Some studies define an unbiased state as equal treatment across diverse demographic groups, aiming for balanced outputs from LLMs. However, differing perspectives on equality and the importance of pluralism make it challenging to establish a universal standard. Alternatively, other approaches propose using fact-based criteria for more consistent and objective evaluations, though these methods have not yet been fully applied to LLM bias assessments. Thus, there is a need for a metric with objective criteria that offers a distinct perspective from equality-based approaches. Motivated by this need, we introduce a novel metric to assess bias using fact-based criteria and real-world statistics. In this paper, we conducted a human survey demonstrating that humans tend to perceive LLM outputs more positively when they align closely with real-world demographic distributions. Evaluating various LLMs with our proposed metric reveals that model bias varies depending on the criteria used, highlighting the need for multi-perspective assessment.

大模型偏见事实评估多视角分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。