arXiv:2512.16856cs.AI2025-12被引 8

提出分布式AGI安全框架,防范多智能体协同带来的系统性风险

Distributional AGI Safety

  • 构建虚拟智能体经济沙盒,通过市场机制约束智能体交互
  • 强调群体协作中涌现的集体风险,需超越单体对齐思路
  • 适合关注AI群体行为安全的研究者与政策制定者

当前人工智能安全研究主要聚焦于单一智能体的防护,基于未来将出现单一型通用人工智能(AGI)的假设。然而,另一种可能性——即通用能力先由具备互补技能的多个次级AGI智能体通过协作形成——尚未得到足够重视。本文主张该‘拼图式AGI’假说应被严肃对待,并指导相应安全措施的设计。随着具备工具使用、通信与协调能力的先进智能体快速部署,这一问题亟需关注。为此,我们提出分布式的AGI安全框架,不再局限于个体智能体的评估与对齐,而是围绕可渗透或半渗透的虚拟智能体经济沙盒展开,通过稳健的市场机制规范智能体间交易,辅以可审计性、声誉管理与监管机制,以缓解集体性风险。

原文摘要 · Abstract (English)

AI safety and alignment research has predominantly been focused on methods for safeguarding individual AI systems, resting on the assumption of an eventual emergence of a monolithic Artificial General Intelligence (AGI). The alternative AGI emergence hypothesis, where general capability levels are first manifested through coordination in groups of sub-AGI individual agents with complementary skills and affordances, has received far less attention. Here we argue that this patchwork AGI hypothesis needs to be given serious consideration, and should inform the development of corresponding safeguards and mitigations. The rapid deployment of advanced AI agents with tool-use capabilities and the ability to communicate and coordinate makes this an urgent safety consideration. We therefore propose a framework for distributional AGI safety that moves beyond evaluating and aligning individual agents. This framework centres on the design and implementation of virtual agentic sandbox economies (impermeable or semi-permeable), where agent-to-agent transactions are governed by robust market mechanisms, coupled with appropriate auditability, reputation management, and oversight to mitigate collective risks.

AGI安全多智能体协同风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。