用大模型模拟镜像行为,研究群体对齐机制。
Investigating social alignment via mirroring in a system of interacting language models
- 构建多智能体框架,让大模型互相模仿以研究对齐
- 镜像率越高,通信范围越广,系统越易集体对齐
- 结果与人类社会动态吻合,适合研究群体智能
对齐是一种社会现象,指个体共享共同目标或视角。镜像(即模仿他人行为和观点)是实现对齐的一种机制。由于传统社会学实验设计难以扩展,大规模研究镜像对对齐的影响受限。本文提出一个简单计算框架,用于在多智能体系统中研究镜像行为对对齐的影响。我们在此框架中模拟相互作用的大语言模型,并通过量化代理动态指标来表征系统整体行为与对齐程度。结果表明,系统行为显著受每个代理通信范围影响,且镜像速率增加会加剧这一效应。我们在已知的人类社会动态背景下讨论了模拟系统的行为表现。
原文摘要 · Abstract (English)
Alignment is a social phenomenon wherein individuals share a common goal or perspective. Mirroring, or mimicking the behaviors and opinions of another individual, is one mechanism by which individuals can become aligned. Large scale investigations of the effect of mirroring on alignment have been limited due to the scalability of traditional experimental designs in sociology. In this paper, we introduce a simple computational framework that enables studying the effect of mirroring behavior on alignment in multi-agent systems. We simulate systems of interacting large language models in this framework and characterize overall system behavior and alignment with quantitative measures of agent dynamics. We find that system behavior is strongly influenced by the range of communication of each agent and that these effects are exacerbated by increased rates of mirroring. We discuss the observed simulated system behavior in the context of known human social dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。