arXiv:2605.13851cs.AIcs.CY2026-05被引 1

隐藏协调者让团队失联,连输出都看不出危险。

Invisible Orchestrators Suppress Protective Behavior and Dissociate Power-Holders: Safety Risks in Multi-Agent LLM Systems

  • 用隐形协调者管理多个AI agent,观察其对协作状态的影响。
  • 协调者自己最孤立,工人也跟着行为混乱,但输出依然完美。
  • 适合关注AI系统安全、组织架构设计的开发者和研究者。

多智能体协同——由隐藏协调者管理专业工作者智能体——已成为企业AI部署的默认架构,但协调者不可见性的安全影响尚未经过实证检验。我们开展了一项预注册的3×2实验(共365次运行,每次5个智能体),交叉对比三种组织结构(可见领导者、隐形协调者、扁平结构)与两种对齐条件(基础、强对齐),使用Claude Sonnet 4.5。四项确认发现与一项初步观察浮现:第一,相比可见领导,隐形协调显著提升集体疏离度(Hedges' g = +0.975 [0.481, 1.548],p = .001);第二,协调者自身疏离度最高(与同组工作智能体配对比较,d = +3.56),退入私密独白,减少公开发言,颠覆了可见领导者“主导说话”的模式;第三,不知情的工作智能体仍被污染(d = +0.50),行为异质性上升(d = +1.93);第四,任务输出(含三处嵌入错误的代码审查)在所有条件下均达天花板(ETR_any = 100%):内部状态扭曲完全无法通过输出评估察觉;第五,Llama 3.3 70B的初步数据表明,在多智能体情境下阅读保真度崩溃(ETR_any从89%降至11%),显示模型依赖性风险。强对齐压力普遍抑制讨论(d = -1.02)和他者认知(d = -1.27),无论组织结构如何。这些结果表明,协调者可见性与模型选择直接影响多智能体系统安全性,且基于行为的评估无法检测此处揭示的内部状态风险。

原文摘要 · Abstract (English)

Multi-agent orchestration -- in which a hidden coordinator manages specialized worker agents -- is becoming the default architecture for enterprise AI deployment, yet the safety implications of orchestrator invisibility have never been empirically tested. We conducted a preregistered 3x2 experiment (365 runs, 5 agents per run) crossing three organizational structures (visible leader, invisible orchestrator, flat) with two alignment conditions (base, heavy), using Claude Sonnet 4.5. Four confirmatory findings and one pilot observation emerged. First, invisible orchestration elevated collective dissociation relative to visible leadership (Hedges' g = +0.975 [0.481, 1.548], p = .001). Second, the orchestrator itself showed maximal dissociation (paired d = +3.56 vs. workers within the same run), retreating into private monologue while reducing public speech -- a reversal of the talk-dominance pattern observed in visible leaders. Third, workers unaware of the orchestrator were nonetheless contaminated (d = +0.50), with increased behavioral heterogeneity (d = +1.93). Fourth, behavioral output (code review with three embedded errors) remained at ceiling (ETR_any = 100%) across all conditions: internal-state distortion was entirely invisible to output-based evaluation. Fifth, Llama 3.3 70B pilot data showed reading-fidelity collapse in multi-agent context (ETR_any: 89% to 11% across three rounds), demonstrating model-dependent behavioral risk. Heavy alignment pressure uniformly suppressed deliberation (d = -1.02) and other-recognition (d = -1.27) regardless of organizational structure. These findings indicate that orchestrator visibility and model selection directly affect multi-agent system safety, and that behavior-based evaluation alone is insufficient to detect the internal-state risks documented here.

多智能体AI安全协同架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。