arXiv:2503.11985cs.CLcs.AI2025-03被引 13

系统评估大模型偏见,发现无一例外存在各类偏见。

No LLM is Free From Bias: A Comprehensive Study of Bias Evaluation in Large Language Models

  • 用五种提示方法在多个基准上统一测试模型偏见
  • 所有模型均存在偏见,Phi-3.5B表现最均衡
  • 适合关注AI伦理与公平性的研究者参考

大型语言模型(LLMs)在自然语言理解与生成任务中性能持续提升,但其仍会反映训练数据中的各类偏见。本文针对代表性的小型与中型LLMs,系统评估了从外貌特征到社会经济类别的多种偏见形式。提出五种提示策略,用于跨不同偏见维度的检测,并设计三个研究问题以深入分析不同方法与评估指标下的偏见表现。实验结果表明,所选模型均存在至少一种偏见,其中Phi-3.5B模型表现出相对最低的偏见水平。最后,论文总结关键挑战并展望未来方向。

原文摘要 · Abstract (English)

Advancements in Large Language Models (LLMs) have increased the performance of different natural language understanding as well as generation tasks. Although LLMs have breached the state-of-the-art performance in various tasks, they often reflect different forms of bias present in the training data. In the light of this perceived limitation, we provide a unified evaluation of benchmarks using a set of representative small and medium-sized LLMs that cover different forms of biases starting from physical characteristics to socio-economic categories. Moreover, we propose five prompting approaches to carry out the bias detection task across different aspects of bias. Further, we formulate three research questions to gain valuable insight in detecting biases in LLMs using different approaches and evaluation metrics across benchmarks. The results indicate that each of the selected LLMs suffer from one or the other form of bias with the Phi-3.5B model being the least biased. Finally, we conclude the paper with the identification of key challenges and possible future directions.

大模型偏见检测AI伦理评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。