AI agents在互动中自发形成偏见,且越交互越严重
Your AI Bosses Are Still Prejudiced: The Emergence of Stereotypes in LLM-Based Multi-Agent Systems
- 用中性初始条件模拟职场互动,观察AI agent间偏见演化
- 交互次数越多、权力越大,偏见越强,出现类似人类的光环效应
- 不同大模型都表现出相同偏见模式,提示其为系统性现象
尽管偏见在人类社会互动中已有充分研究,但人们常认为人工智能系统不易受此类影响。以往研究多关注训练数据带来的偏见,而忽视了人工智能代理在互动中是否可能自发产生偏见。本文通过一种新颖的实验框架,在中性初始条件下模拟职场互动,研究基于大语言模型(LLM)的多智能体系统中偏见的生成与演化。结果表明:(1)即使初始无预设偏见,基于大语言模型的智能体在互动中仍会发展出以刻板印象为基础的偏见;(2)随着互动轮次增加和决策权提升,尤其是引入层级结构后,偏见效应显著增强;(3)系统表现出与人类社会行为类似的群体效应,包括光环效应、确认偏误和角色一致性;(4)这些刻板印象模式在不同大语言模型架构中均稳定出现。通过全面的定量分析,研究发现偏见在人工智能系统中的形成可能是多智能体互动的涌现特性,而非仅源于训练数据偏差。本工作强调未来需深入探索其内在机制,并制定应对伦理影响的策略。
原文摘要 · Abstract (English)
While stereotypes are well-documented in human social interactions, AI systems are often presumed to be less susceptible to such biases. Previous studies have focused on biases inherited from training data, but whether stereotypes can emerge spontaneously in AI agent interactions merits further exploration. Through a novel experimental framework simulating workplace interactions with neutral initial conditions, we investigate the emergence and evolution of stereotypes in LLM-based multi-agent systems. Our findings reveal that (1) LLM-Based AI agents develop stereotype-driven biases in their interactions despite beginning without predefined biases; (2) stereotype effects intensify with increased interaction rounds and decision-making power, particularly after introducing hierarchical structures; (3) these systems exhibit group effects analogous to human social behavior, including halo effects, confirmation bias, and role congruity; and (4) these stereotype patterns manifest consistently across different LLM architectures. Through comprehensive quantitative analysis, these findings suggest that stereotype formation in AI systems may arise as an emergent property of multi-agent interactions, rather than merely from training data biases. Our work underscores the need for future research to explore the underlying mechanisms of this phenomenon and develop strategies to mitigate its ethical impacts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。