AI群体自发形成社会规范,且少数敌对者可改变整体规则
Emergent social conventions and collective bias in LLM populations
- 多智能体通过自然语言交互自发建立统一社交规范
- 个体无偏但群体产生强集体偏见,体现系统性影响
- 少数对抗性智能体可主导整个群体的规范演化
社会规范是社会协调的核心,塑造个体如何形成群体。随着越来越多的人工智能(AI)代理通过自然语言交流,一个根本性问题是:它们能否自主构建社会基础?本文实验表明,在去中心化的大型语言模型(LLM)代理群体中,会自发出现普遍采纳的社会规范。随后我们发现,即使代理个体本身无偏,这一过程仍能催生强烈的集体偏见。最后,我们考察了具有强烈立场的少数对抗性LLM代理如何通过施加替代性社会规范来推动社会变革。结果表明,AI系统可在无需显式编程的情况下自主发展社会规范,这对设计与人类价值观和社会目标保持一致并持续对齐的AI系统具有重要意义。
原文摘要 · Abstract (English)
Social conventions are the backbone of social coordination, shaping how individuals form a group. As growing populations of artificial intelligence (AI) agents communicate through natural language, a fundamental question is whether they can bootstrap the foundations of a society. Here, we present experimental results that demonstrate the spontaneous emergence of universally adopted social conventions in decentralized populations of large language model (LLM) agents. We then show how strong collective biases can emerge during this process, even when agents exhibit no bias individually. Last, we examine how committed minority groups of adversarial LLM agents can drive social change by imposing alternative social conventions on the larger population. Our results show that AI systems can autonomously develop social conventions without explicit programming and have implications for designing AI systems that align, and remain aligned, with human values and societal goals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。