超50个大模型实测发现:模型越大越复杂,隐性偏见可能越严重。
Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs
- 基于IAT和决策偏见框架,系统检测50+大模型的隐性偏见
- 新模型如Llama、GPT系列在部分任务中反而比旧版偏见更高
- 揭示行业缺乏统一评估标准,适合关注AI伦理的研究者与开发者
大型语言模型(LLMs)被广泛应用于各类任务,包括对公平性要求极高的决策场景。尽管通过显式偏见测试,部分模型仍存在隐性偏见。本研究基于LLM隐性关联测试(IAT Bias)与LLM决策偏见框架,对超过50个主流大模型进行大规模评估,发现更先进或更大规模的模型并未自动降低偏见,甚至在某些情况下(如Meta的Llama系列和OpenAI的GPT模型)表现出比前代更高的偏见分数。这表明,在缺乏明确偏见缓解策略的情况下,模型复杂度提升可能无意中放大既有偏见。模型间及内部偏见评分的显著差异,凸显了建立标准化评估指标与基准的迫切需求。当前偏见缓解尚未成为普遍优先目标,可能导致不公平或歧视性结果。该研究通过扩展隐性偏见检测范围,为理解先进模型中的偏见提供了更全面视角,并强调了构建公平、负责任AI系统的重要性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are being adopted across a wide range of tasks, including decision-making processes in industries where bias in AI systems is a significant concern. Recent research indicates that LLMs can harbor implicit biases even when they pass explicit bias evaluations. Building upon the frameworks of the LLM Implicit Association Test (IAT) Bias and LLM Decision Bias, this study highlights that newer or larger language models do not automatically exhibit reduced bias; in some cases, they displayed higher bias scores than their predecessors, such as in Meta's Llama series and OpenAI's GPT models. This suggests that increasing model complexity without deliberate bias mitigation strategies can unintentionally amplify existing biases. The variability in bias scores within and across providers underscores the need for standardized evaluation metrics and benchmarks for bias assessment. The lack of consistency indicates that bias mitigation is not yet a universally prioritized goal in model development, which can lead to unfair or discriminatory outcomes. By broadening the detection of implicit bias, this research provides a more comprehensive understanding of the biases present in advanced models and underscores the critical importance of addressing these issues to ensure the development of fair and responsible AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。