多智能体安全不取决于模型大小,而由交互结构决定。
Position: Safety and Fairness in Agentic AI Depend on Interaction Topology, Not on Model Scale or Alignment

- 交互拓扑决定安全:顺序决策或投票机制影响整体行为
- 三种病理现象:顺序敏感、信息级联、公平性形式化但风险识别失效
- 模型越大越危险,需从系统层面评估交互结构
随着大语言模型作为智能体参与高风险决策,学界普遍认为个体模型的安全性可组合为多智能体系统的安全性。本文指出该假设根本错误。在智能体人工智能中,安全由交互拓扑决定,而非模型权重。当智能体按序讨论或通过并行投票+裁判聚合时,信息流与决策耦合结构主导结果。跨模型家族与规模的证据显示三种持续存在的拓扑驱动病态:顺序不稳定性(系统行为依赖智能体顺序)、信息级联(早期判断无论对错均传播)、功能坍缩(满足公平指标但放弃真实风险判别)。反直觉的是,模型能力增强会强化这些效应,因共识形成更强且初始决策更难挑战。这些故障模式无法被以模型为中心的评估与对齐方法发现。我们主张将智能体AI视为动态系统,而非对齐组件的集合。交互拓扑必须成为安全评估与监管的核心,部署前需验证系统在不同架构下的鲁棒性。
原文摘要 · Abstract (English)
As large language models are increasingly deployed as interacting agents in high-stakes decisions, the AI safety community assumes that safety properties of individual models will compose into safe multi-agent behavior. This position paper argues that this assumption is fundamentally mistaken. In agentic AI, safety is determined by interaction topology, not model weights. When agents deliberate sequentially or aggregate via parallel voting with a judge, the structure of information flow and decision coupling dominates outcomes. Evidence across model families and scales reveals three persistent topology-driven pathologies: ordering instability, where system behavior depends primarily on agent sequence; information cascades, where early judgments propagate regardless of correctness; and functional collapse, where systems satisfy fairness metrics while abandoning meaningful risk discrimination. Contrary to intuition, scaling to more capable models strengthens these effects by increasing consensus formation and reducing the challenge of initial decisions. These failure modes are invisible to model-centric evaluation and alignment procedures. We argue that agentic AI must be treated as a dynamical system rather than a collection of aligned components. Interaction topology must become a primary target of safety evaluation and regulation, with systems required to demonstrate robustness across architectural variations before deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。