大模型会无意识模仿人类行为中的偏见,即使这些行为与真实偏好无关。
Mimicry without understanding: the origins of decision bias in large language models
- 大模型会错误模仿人类行为中的偏见,哪怕行为与偏好无逻辑关联。
- 当报告描述损失厌恶时,模型自身也表现出相似偏见,且程度与报告一致。
- 研究揭示了大模型偏见的生成机制,适合关注AI伦理与安全的研究者。
大型语言模型(LLMs)易受社会、情感和认知偏见影响。本文探讨了两种即使在训练数据中无偏或已正确标注为偏见的情况下仍会产生偏见的机制:一是基于人类行为的错误模仿,即模型在行为与偏好无逻辑关联时仍推断出偏好;二是对显性偏见行为的模仿。四项聚焦经济偏见的研究发现,当提示中包含明显不能反映真实偏好的人类行为报告时,ChatGPT-4o 和 Qwen 仍表现出社会证明偏见。此外,当明确描述损失厌恶这一偏见时,模型自身也表现出该偏见。更关键的是,在提示中提供详细科学报告后,模型的偏见程度与报告中描述的偏见强度高度相关。这表明,关于偏见的科学文献可能成为大模型的自我实现预言。
原文摘要 · Abstract (English)
Large Language models (LLMs) were found to be susceptible to a host of social, affective, and cognitive biases. We examined two mechanisms through which such biases can be generated even when human preferences (in the training data) are not biased or when they are correctly categorized as being biased. The first is faulty mimicry of preferences based on human behavior: this involves LLMs inferring human preferences even when behaviors are logically unrelated to preferences. The second is mimicry of explicitly biased human behaviors. In four studies focusing on economic biases, we find that ChatGPT-4o and Qwen exhibited social proof biases even when prompted with reports of human behaviors that were clearly non-indicative of individuals' actual preferences. LLMs also displayed loss aversion when it was explicitly described as a bias. Indeed, when prompted with detailed scientific reports, the extent of the bias (i.e., loss aversion) in the scientific report predicted LLMs' own subsequent bias. Scientific papers of biases can thus become self-fulfilling prophecies, at least when it comes to LLMs' responses. The current study goes beyond fleshing out LLM biases and sheds light on the underlying component processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。