用统计物理方法分析大模型多智能体系统,区分内在偏见与合作行为。
Collective Alignment in LLM Multi-Agent Systems: Disentangling Bias from Cooperation via Statistical Physics

- 构建二维网格大模型智能体系统,通过温度调控研究集体动态。
- 发现各模型均呈现有序-无序转变,但未达伊辛模型临界指数。
- 提取有效耦合与场参数,揭示内在偏见主导集体对齐现象。
我们研究基于大语言模型的多智能体系统在二维方格晶格上的涌现集体动力学,提出一种模型无关的统计物理方法,用于分离社会从众与内在偏见,计算临界指数,并探测系统的集体行为及可能的相变。在框架中,每个 $L\times L$ 网格节点部署一个相同的LLM智能体,状态为二元(+1/-1,对应是/否),根据四个最近邻状态查询模型更新。采样温度 $T$ 作为唯一控制参数。在三个开源模型(llama3.1:8b, phi4-mini:3.8b, mistral:7b)上,通过全局翻转协议测量磁化率与磁化强度。所有模型均显示温度驱动的有序-无序交叉,且出现峰值;在偶数-$L$ 网格上进行有限尺寸标度分析,得到模型依赖的有效指数 $γ/ν$,其值接近但不兼容二维伊辛普适类($γ/ν=7/4$)。该方法可提取有效 $β$-加权耦合 $ ilde{J}(T)$ 与场 $ ilde{h}(T)$,分别表征社会从众与内在偏见。分析表明,集体对齐主要由内在偏见($ ilde{h}\gg\tilde{J}$)主导,而非邻居协作耦合,导致场驱动的交叉而非真实相变。这些有效参数在不同模型间定性差异显著,为大模型智能体提供紧凑的集体行为指纹,并可用于量化多智能体共识与对齐的可靠性诊断。
原文摘要 · Abstract (English)
We investigate the emergent collective dynamics of LLM-based multi-agent systems on a 2D square lattice and present a model-agnostic statistical-physics method to disentangle social conformity from intrinsic bias, compute critical exponents, and probe the collective behavior and possible phase transitions of multi-agent systems. In our framework, each node of an $L\!\times\!L$ lattice hosts an identical LLM agent holding a binary state ($+1$/$-1$, mapped to yes/no) and updating it by querying the model conditioned on the four nearest-neighbor states. The sampler temperature $T$ serves as the sole control parameter. Across three open-weight models (llama3.1:8b, phi4-mini:3.8b, mistral:7b), we measure magnetization and susceptibility under a global-flip protocol designed to probe $\mathbb{Z}_2$ symmetry. All models display temperature-driven order-disorder crossovers and susceptibility peaks; finite-size scaling on even-$L$ lattices yields effective exponents $γ/ν$ whose values are model-dependent, close to but incompatible with the 2D Ising universality class ($γ/ν=7/4$). Our method enables the extraction of effective $β$-weighted couplings $\tilde{J}(T)$ and fields $\tilde{h}(T)$, which serve as a measure of social conformity and intrinsic bias. In the models we analyzed, we found that collective alignment is dominated by an intrinsic bias ($\tilde{h}\gg\tilde{J}$) rather than by cooperative neighbor coupling, producing field-driven crossovers instead of genuine phase transitions. These effective parameters vary qualitatively across models, providing compact collective-behavior fingerprints for LLM agents and a quantitative diagnostic for the reliability of multi-agent consensus and collective alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。