arXiv:2511.06148cs.CYcs.AI2025-11被引 4

大模型在决策中会自发产生新型社会偏见,且越大的模型越严重。

Large Language Models Develop Novel Social Biases Through Adaptive Exploration

  • 通过模拟探索与利用的权衡,模型在无差异群体间形成偏见
  • 新模型任务分配更不公,公平性低于人类参与者
  • 明确激励探索可有效降低偏见,适合关注模型伦理的研究者

随着大语言模型(LLMs)被应用于具备真实决策能力的框架,确保其无偏变得愈发重要。本文指出,单纯消除已有偏见的方法不足以应对问题。基于心理学中的范式,我们发现:即使在不存在本质差异的人工人口群体中,LLMs仍会自发形成新的社会偏见。这些偏见导致高度分层的任务分配,其公平性低于人类参与者的分配结果,且在更新、更大的模型中更为严重。人类中的类似偏见源于探索-利用权衡——决策者探索不足,使早期观察过度影响对整个群体的认知。为此,我们尝试了多种干预措施,包括调整输入、问题结构和显式引导。多数干预效果有限,但显式激励探索能显著减少分层现象,凸显出需要更复杂的多维度目标来缓解偏见。研究揭示,LLMs并非被动反映人类偏见,而是能从经验中主动创造新偏见,引发关于其长期社会影响的紧迫思考。

原文摘要 · Abstract (English)

As large language models (LLMs) are adopted into frameworks that grant them the capacity to make real decisions, it is increasingly important to ensure that they are unbiased. In this paper, we argue that the predominant approach of simply removing existing biases from models is not enough. Using a paradigm from the psychology literature, we demonstrate that LLMs can spontaneously develop novel social biases about artificial demographic groups even when no inherent differences exist. These biases result in highly stratified task allocations, which are less fair than assignments by human participants and are exacerbated in newer and larger models. In humans, emergent biases like these have been shown to result from exploration-exploitation trade-offs, where the decision-maker explores too little, allowing early observations to strongly influence impressions about entire demographic groups. To alleviate this effect, we explore a series of interventions targeting model inputs, problem structure, and explicit steering. While most interventions have limited effect, explicitly incentivizing exploration robustly reduces stratification, highlighting the need for better multifaceted objectives to mitigate bias. These results reveal that LLMs are not merely passive mirrors of human social biases, but can actively create new ones from experience, raising urgent questions about how these systems will shape societies over time.

大模型社会偏见探索利用公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。