arXiv:2501.17420cs.CLcs.AI2025-01被引 38

用虚拟角色决策测试发现大模型隐性偏见,越先进的模型越明显。

Actions Speak Louder than Words: Agent Decisions Reveal Implicit Biases in Language Models

  • 让大模型扮演不同社会背景的虚拟角色,观察其决策差异。
  • 六款主流模型在多数场景中均显显著社会群体偏差。
  • 结果与真实世界偏差方向一致但更严重,适合评估模型公平性。

尽管公平性与对齐技术已缓解大语言模型在显式提示下的外显偏见,我们推测这些模型在模拟人类行为时仍可能表现出隐性偏见。为此,我们提出一种系统方法,通过评估具有社会人口学特征的虚拟角色(由大模型生成)在决策中的差异,揭示跨广泛社会人口类别中的隐性偏见。我们在三个社会人口组和四个决策场景下测试了六款大语言模型。结果表明,当前最先进的大模型在几乎所有模拟中均存在显著的社会人口差异,且先进模型虽减少了显性偏见,却表现出更强的隐性偏见。与实证研究报道的真实世界偏差相比,我们发现的偏见方向一致但显著放大。这一方向一致性凸显了该方法在识别系统性偏见而非随机波动方面的有效性;同时,隐性偏见的存在及其放大,凸显了需采用新策略应对此类问题。

原文摘要 · Abstract (English)

While advances in fairness and alignment have helped mitigate overt biases exhibited by large language models (LLMs) when explicitly prompted, we hypothesize that these models may still exhibit implicit biases when simulating human behavior. To test this hypothesis, we propose a technique to systematically uncover such biases across a broad range of sociodemographic categories by assessing decision-making disparities among agents with LLM-generated, sociodemographically-informed personas. Using our technique, we tested six LLMs across three sociodemographic groups and four decision-making scenarios. Our results show that state-of-the-art LLMs exhibit significant sociodemographic disparities in nearly all simulations, with more advanced models exhibiting greater implicit biases despite reducing explicit biases. Furthermore, when comparing our findings to real-world disparities reported in empirical studies, we find that the biases we uncovered are directionally aligned but markedly amplified. This directional alignment highlights the utility of our technique in uncovering systematic biases in LLMs rather than random variations; moreover, the presence and amplification of implicit biases emphasizes the need for novel strategies to address these biases.

大模型偏见公平性虚拟角色隐性偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。