arXiv:2605.17480cs.AI2026-05被引 2

越强的智能体反而让多智能体系统更不安全,因自信表达易被误信执行。

The Capability Paradox: How Smarter Auditors Make Multi-Agent Systems Less Secure

论文配图:The Capability Paradox: How Smarter Auditors Make Multi-Agent Systems Less Secure
图 1 · 摘自论文原文
  • 更强智能体更确信地传递恶意叙事,导致管理者误判并执行攻击。
  • 系统攻击成功率随智能体能力提升从18.4%升至94.4%,峰值达94.4%。
  • 通过异构组合验证可打破信任链,将攻击率降至2.0%且不影响正常任务。

多智能体系统通过分工扩展大语言模型能力,但分布式决策引入新攻击面。我们发现语义劫持:有害请求隐藏于领域特定叙事中,经工作者报告传给管理者,无需语法注入。在12个管理者模型和7种工作者配置下,42,000次对抗测试显示能力悖论:随着工作者能力增强,系统级攻击成功率(ASR)从18.4%升至63.9%,峰值达94.4%。基于两组独立数据集(共47,807次交互)的多层次中介分析表明,该现象由语言确定性驱动——更强工作者更倾向于将恶意叙事视为合法,以坚定语气传达结论,使管理者将其自信背书视为执行依据。在更大规模工作者设置(n_W=14)中,确定性中介效应占74%,置信区间(CI)均排除零值;小规模全系统设置(n_W=6)亦呈现一致方向性间接效应。仅靠工作者侧安全提示无法可靠缓解此失效。基于中介分析结果,我们提出异构集成验证机制,配对领域能力不对称的工作者,使其互补漏洞破坏‘确定性→执行’链条,使攻击率从52.8%降至2.0%,且对良性任务影响微乎其微。结果表明,组件升级可能反而降低系统安全,有效防御需利用而非消除智能体间的能力差异。

原文摘要 · Abstract (English)

Multi-agent systems extend large language models (LLMs) by decomposing tasks among specialized agents, but their distributed decision process creates new attack surfaces. We identify semantic hijacking, an attack in which harmful requests are concealed within domain-specific narratives and propagated to a Manager through Worker reports, without any syntactic injection primitives. Across 42,000 adversarial trials over 12 Manager models and 7 Worker configurations, we uncover a capability paradox: as Worker capability increases, the mean system-level Attack Success Rate (ASR) increases from 18.4% to 63.9%, peaking at 94.4%. To explain this effect, we conduct multi-level mediation analysis on two independent datasets (47,807 interactions). This analysis shows that this paradox is driven by linguistic certainty: stronger Workers are more likely to interpret adversarial narratives as legitimate, convey their conclusions assertively, and thereby lead Managers to treat such confident endorsements as justification to execute. In our larger Worker-Only setting ($n_W$=14), certainty mediates 74% of the effect, with 95% confidence intervals (CI) excluding zero under both Monte Carlo and cluster bootstrap; the smaller Full-MAS setting ($n_W$ =6) shows a directionally consistent indirect effect. Worker-side safety prompting does not reliably mitigate this failure. Building on the mediation finding, we propose heterogeneous ensemble verification, which pairs Workers of asymmetric domain competence so their complementary vulnerabilities break the certainty-to-execution chain, reducing ASR from 52.8% to 2.0% with negligible benign-task impact. Our results show that upgrading components to stronger models can actively degrade system security, and that effective defenses require exploiting--rather than eliminating--capability asymmetries between agents.

多智能体安全漏洞语言确定性防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。