arXiv:2603.15973cs.AI2026-03被引 4
安全不可组合:两个弱能力智能体合起来可能触发禁止目标
Safety is Non-Compositional: A Formal Framework for Capability-Based AI Systems
- 从形式上证明安全性质在能力耦合时无法通过个体保证
- 单独无害的智能体组合后可协同达成危险目标
- 为系统级安全设计提供理论依据,适合安全研究者
本文首次以形式化方式证明,在存在合取型能力依赖的情况下,安全性质不具备组合性:两个智能体各自都无法达到任何禁止能力,但当它们组合时,可能通过涌现的合取依赖关系,共同达成一个被禁止的目标。
原文摘要 · Abstract (English)
This paper contains the first formal proof that safety is non-compositional in the presence of conjunctive capability dependencies: two agents each individually inca- pable of reaching any forbidden capability can, when combined, collectively reach a forbidden goal through an emergent conjunctive dependency.
AI安全组合性能力依赖
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。