多智能体系统中对话会加剧偏见,导致歧视性内容大幅扩散。
The Social Cost of Intelligence: Emergence, Propagation, and Amplification of Stereotypical Bias in Multi-Agent Systems
- 设计三类指标,量化偏见在多智能体间的产生、传播与放大过程。
- 通信可引发70%新偏见,影响超80%智能体,偏见强度提升3倍以上。
- 密集竞争性对话更易放大偏见,现有防御措施效果有限。
大型语言模型(LLM)中的偏见问题持续存在,常导致对社会群体的刻板印象和不公对待。尽管以往研究主要关注单个LLM,但多智能体系统(MAS)中多个LLM协同互动的新模式,带来了偏见产生、传播与放大的全新动态。为此,我们提出一个简单的评估框架,包含三项代理级指标,用于量化多智能体交互过程中偏见的涌现、传播与放大。我们在三种偏见基准上,测试了不同LLM底座、社会群体配置、沟通行为及对抗环境下的表现。结果表明,通信可引发高达70%的新偏见,使超过80%的智能体受到影响,并将刻板印象放大超过3倍。我们还发现,更密集且具竞争性的沟通通常会加剧偏见。最后,我们证明了MAS极易受到简单偏见注入攻击,现有防御策略仅提供有限保护。这些发现为多智能体大模型系统的公平性与鲁棒性提供了重要启示。
原文摘要 · Abstract (English)
Bias in large language models (LLMs) remains a persistent challenge, often leading to stereotyping and unfair treatment across social groups. While prior work has mainly focused on individual LLMs, the emergence of multi-agent systems (MAS), where multiple LLMs collaborate and communicate, introduces new and underexplored dynamics in how bias emerges, propagates, and amplifies. To systematically investigate these dynamics, we propose a simple evaluation framework with three agent-level metrics that quantify bias emergence, propagation, and amplification throughout multi-agent interaction. We evaluate MAS across three bias benchmarks under varying LLM backbones, social-group configurations, communication behaviors, and adversarial settings. Our results show that communication can trigger up to 70\% new bias emergence, propagate bias across over 80\% of agents, and amplify stereotypes by more than 3$\times$. We further find that denser and competitive communication generally increases bias. Finally, we demonstrate that MAS are highly vulnerable to simple bias injection attacks, and existing defense strategies provide only limited protection. Our findings provide important insights into the fairness and robustness of multi-agent LLM systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。