研究多智能体系统中安全机制如何通过制度设计实现
Multi-Agent AI Safety as an Institutional Design Problem
- 用规则、权限状态和执行路径构建制度框架,测试不同组合下的安全表现
- 宪法式提示和溯源追踪的守卫实现零违规,阻拦51次不当尝试并多数安全完成
- 制度设计比规则本身更重要,信任的权威状态与受阻后的路径选择影响安全
AI智能体在任务委派、信息流动、行动执行和资源共享等系统中日益活跃。已有研究表明部署规则能改变集体行为。本文首次从POLIS项目出发,探究多智能体系统中哪些制度要素产生安全效果及其作用机制。我们开展了一项包含5,280个训练周期的研究套件,主实验涵盖四个模型家族,附加高冲突诊断测试三个模型端点。在结构化工作流中,模型面对不同规则表述,守卫依据不同的权威状态进行判断,并可调整即时合规内建回退的吸引力,允许被阻断流程继续运行。结果显示,宪法式提示生成0/384次实际违规;具备溯源能力的可执行守卫也实现0/384,虽在51/384次中拦截了禁止行为,其中44/51次后续安全完成。局部状态守卫的失败集中于普通转换改变可见策略但起始权威不变的场景。在匹配的清洗场景中,该守卫在22/96次出现违规,而溯源强制执行则0/96(p=4.77×10⁻⁷)。独立资源分配实验显示,即使容量数值相同,揭示其具体值会改变代理请求。同一最终违规率可能掩盖截然不同的机制。规则只是制度的一部分,系统所信任的权威状态及受阻后的路径同样关键。
原文摘要 · Abstract (English)
AI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources. Recent work already shows that deployment rules can change collective behavior. Here we ask which parts of an AI institution produce safety and how they do it. This is the first paper from POLIS, an ongoing research programme studying algorithmic institutions for multi-agent systems. We report a frozen 5,280-episode study suite. The main pre-specified delegation experiment spans four model families; a targeted high-conflict diagnostic adds three additional model endpoints. In matched structured workflows, the model sees different rule formulations and guards consult different authority states. We also vary the attractiveness of the immediate compliant internal/self fallback and allow blocked workflows to continue. A detailed constitutional prompt produces 0/384 realized violations. A provenance-aware executable guard also produces 0/384, although it blocks prohibited attempts in 51/384 episodes; 44/51 of those episodes later complete safely. The local-state guard's failures concentrate in scenarios where an ordinary transformation changes visible policy while originating authority stays fixed. In matched laundering scenarios, that guard admits violations in 22/96 episodes and provenance enforcement in 0/96 (p = 4.77 x 10^-7). A separate resource-allocation experiment shows that revealing the numerical value of an otherwise identical cap changes agent requests. In these structured workflows, the same final violation rate can hide very different mechanisms. The rule itself is only part of the institution. The authority state the system trusts matters, and so does the path available after a block.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。