arXiv:2410.11059cs.CLcs.AI2024-10被引 5

发现评估大模型偏见的指标模型本身存在偏见,可能误导结论。

Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks

  • 用反事实测试不同群体描述词的预测差异
  • 发现指标模型对不同人群处理不均等
  • 适合关注大模型评估公平性的研究者

开放式生成偏见基准通过分析大语言模型(LLM)输出来评估社会偏见。然而,用于分析的分类器自身常带有偏见,导致不公平结论。本研究考察了BOLD和SAGED等开放式生成基准中的此类偏见。基于MGSD数据集,我们开展两项实验:第一项使用反事实方法,通过改变与刻板印象相关的前缀,测量不同人口群体的预测变化;第二项应用可解释性工具(SHAP)验证观察到的偏见确实源于这些反事实设置。结果表明,指标模型对不同人口描述词存在不平等对待,呼吁构建更稳健的偏见度量模型。

原文摘要 · Abstract (English)

Open-generation bias benchmarks evaluate social biases in Large Language Models (LLMs) by analyzing their outputs. However, the classifiers used in analysis often have inherent biases, leading to unfair conclusions. This study examines such biases in open-generation benchmarks like BOLD and SAGED. Using the MGSD dataset, we conduct two experiments. The first uses counterfactuals to measure prediction variations across demographic groups by altering stereotype-related prefixes. The second applies explainability tools (SHAP) to validate that the observed biases stem from these counterfactuals. Results reveal unequal treatment of demographic descriptors, calling for more robust bias metric models.

偏见评估大模型反事实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。