提出可扩展至百万级智能体的归因方法,揭示小样本研究会严重扭曲群体行为的根源。
Attributing Emergence in Million-Agent Systems

- 将数学归因理论改造为适用于百万智能体系统的新算法
- 在真实社交数据中发现小样本与全规模结果结构差异显著
- 证明非线性指标下无法通过事后调整弥补偏差,全规模分析不可替代
大语言模型(LLMs)可模拟个体智能体的人类式推理与决策。基于LLM的多智能体系统(MAS)通过组合这些智能体,模拟大规模社会现象如极化、信息传播和市场恐慌。这类研究需将宏观涌现现象归因于个体智能体,但现有公理化方法随智能体数量N呈组合爆炸增长,仅限于N≤10³,而实际现象发生在N≥10⁶量级。本文通过将Aumann–Shapley路径积分归因法适配至百万级LLM-MAS,使方法满足全部四条公理,运行速度比采样Shapley快3至5个数量级,并将可行公理化归因规模提升超过三个数量级(达1670倍)。我们以14天的Bluesky公开数据(1,671,587活跃用户,五个话题)实证验证:全规模归因与常用于小规模研究的可视性偏差样本(N=10²)结果在结构上存在根本分歧——全规模下长尾与中层智能体共同承担主要影响,而小样本将约两倍的影响份额错误归于高层关注者(48%对比24%)。进一步证明,这种分歧无法通过事后全局缩放修复:只有当宏观指标对智能体呈线性关系时,才存在统一缩放因子;我们的非线性指标残差为0.10–0.98。因此,对于非线性指标,全规模归因是必要要求而非选择。
原文摘要 · Abstract (English)
Large language models (LLMs) can simulate human-like reasoning and decision-making in individual agents. LLM-powered multi-agent systems (MAS) combine such agents to simulate population-scale social phenomena such as polarization, information cascades, and market panics. Such studies require attributing macro emergence to individual agents, but existing axiomatic methods scale combinatorially in $N$ and have been confined to $N \lesssim 10^3$, while the phenomena they explain occur at $N \geq 10^6$. We address this gap by adapting Aumann--Shapley path-integral attribution to LLM-powered MAS at million-agent scale; the resulting method satisfies all four axioms, runs three to five orders of magnitude faster than sampled Shapley on the same hardware, and extends feasible axiomatic attribution by over three orders of magnitude (a $1670\times$ jump). We use this method to test the scale gap empirically: across 14 days of public Bluesky data ($1{,}671{,}587$ active users, five topics), we compute the attribution at both full scale and the visibility-biased $N = 10^2$ convenience sample used by small-scale studies, and the two disagree structurally. At full scale the long tail and middle tier jointly carry the majority; the biased small panel shifts about twice that share onto the upper follower tiers ($48\%$ versus $24\%$). We then prove that the disagreement cannot in general be reduced by post-hoc rescaling: an Attribution Scaling Bias theorem shows that a reconciling global rescaling factor exists exactly when the macro indicator is linear over agents, and our nonlinear indicators give residuals of $0.10$--$0.98$. For such nonlinear indicators, full-scale attribution is therefore a requirement rather than a methodological choice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。