arXiv:2605.00248cs.AIcs.GT2026-05

用因果模型判断多个智能体是否形成有共同目标的集体智能。

Causal Foundations of Collective Agency

  • 通过因果博弈与抽象理论,从行为合理性判断集体智能是否存在。
  • 揭示了演员-评论家模型中多智能体激励的矛盾现象。
  • 可量化评估投票机制等场景下的集体智能程度,适合安全研究者参考。

先进AI系统的一个关键安全挑战是:多个简单智能体可能无意中形成一个具有独立能力与目标的集体智能体。更普遍地,如何判定一组智能体能否被视为统一的集体智能体,是生物与人工系统交互和激励研究的基础问题。本文从行为视角出发,当一组智能体的联合行为可被理性且目标导向地预测时,即认为其具备集体智能。我们利用因果博弈(causal games)——对多智能体策略互动的因果建模——与因果抽象(causal abstraction)——形式化高阶模型如何忠实捕捉低阶模型——来构建该框架。该方法解决了演员-评论家模型中多智能体激励的悖论,并对不同投票机制的集体智能程度进行了定量评估。本框架旨在为理解、预测与控制多智能体系统中涌现的集体智能提供理论与实证基础。

原文摘要 · Abstract (English)

A key challenge for the safety of advanced AI systems is the possibility that multiple simpler agents might inadvertently form a collective agent with capabilities and goals distinct from those of any individual. More generally, determining when a group of agents can be viewed as a unified collective agent is a foundational question in the study of interactions and incentives in both biological and artificial systems. We adopt a behavioral perspective in answering this question, ascribing collective agency to a group when viewing the group's joint actions as rational and goal-directed successfully predicts its behavior. We formalize this perspective on collective agency using causal games -- which are causal models of strategic, multi-agent interactions -- and causal abstraction -- which formalizes when a simple, high-level model faithfully captures a more complex, low-level model. We use this framework to solve a puzzle regarding multi-agent incentives in actor-critic models and to make quantitative assessments of the degree of collective agency exhibited by different voting mechanisms. Our framework aims to provide a foundation for theoretical and empirical work to understand, predict, and control emergent collective agents in multi-agent AI systems.

集体智能因果建模多智能体AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。