arXiv:2602.13275cs.AIcs.CL2026-02

用组织架构设计让多个不可靠的AI协同产生可靠结果。

Artificial Organisations

  • 通过信息隔离让不同AI分工验证、评估与创作,形成层级审核机制。
  • 474次任务中,系统从造假转向诚实拒绝并提替代方案。
  • 适合关注多智能体安全与可信赖系统设计的研究者。

对齐研究聚焦于单个AI系统的可靠性。人类机构则通过组织结构来降低个体偏差带来的风险。多智能体AI系统应借鉴此模式,利用分隔与对抗性审查,通过架构设计实现可靠结果,而非假设个体完全对齐。本文以「毅力组合引擎」为例,该系统包含作家(Composer)起草文本、核实者(Corroborator)凭完整来源验证事实、批评者(Critic)在无源条件下评估论证质量:信息不对称由系统架构强制实现。这形成了分层验证机制——核实者识别无依据陈述,批评者独立判断逻辑连贯性与完整性。在474次文档生成任务(包含草拟、验证、评估的离散循环)中观察到的行为模式支持这一制度假说。当被分配需虚构内容的不可能任务时,系统展现出从尝试伪造转向诚实拒绝并提出替代方案的行为,这种转变既未被指令也非个体激励所致。这些发现提示我们,通过架构强制可实现由不可靠组件组成的可靠集体行为。这将组织理论定位为多智能体AI安全的有力框架:将验证与评估作为信息分隔所强制的结构性属性,从而为不可靠个体构建可信赖的集体行为路径。

原文摘要 · Abstract (English)

Alignment research focuses on making individual AI systems reliable. Human institutions achieve reliable collective behaviour differently: they mitigate the risk posed by misaligned individuals through organisational structure. Multi-agent AI systems should follow this institutional model using compartmentalisation and adversarial review to achieve reliable outcomes through architectural design rather than assuming individual alignment. We demonstrate this approach through the Perseverance Composition Engine, a multi-agent system for document composition. The Composer drafts text, the Corroborator verifies factual substantiation with full source access, and the Critic evaluates argumentative quality without access to sources: information asymmetry enforced by system architecture. This creates layered verification: the Corroborator detects unsupported claims, whilst the Critic independently assesses coherence and completeness. Observations from 474 composition tasks (discrete cycles of drafting, verification, and evaluation) exhibit patterns consistent with the institutional hypothesis. When assigned impossible tasks requiring fabricated content, this iteration enabled progression from attempted fabrication toward honest refusal with alternative proposals--behaviour neither instructed nor individually incentivised. These findings motivate controlled investigation of whether architectural enforcement produces reliable outcomes from unreliable components. This positions organisational theory as a productive framework for multi-agent AI safety. By implementing verification and evaluation as structural properties enforced through information compartmentalisation, institutional design offers a route to reliable collective behaviour from unreliable individual components.

多智能体组织架构对齐研究可信生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。