arXiv:2608.08131cs.CRcs.AI2026-08

揭示大模型代理系统中潜伏威胁的组合攻击机制,类比《星战》'66号令'。

Compositional Threat Analysis of Latent Compromise in LLM Agent Systems: The Order 66 Scenario

  • 提出组合式安全模型,分析多组件协同触发破坏性行为的路径。
  • 实证显示跨类别反馈可维持传播,即使同类繁殖率低于1。
  • 适合关注AI系统安全、防御设计的研究者与开发者阅读。

在虚构的‘66号令’中,灾难并非由单一强大指令引发:信任群体被预先设局,短指令激活隐藏条件,而保护机制反而转为攻击。本文将该机制转化为对使用工具的大语言模型(LLM)代理系统的通用安全分析。典型场景包括:部署时植入的潜在破坏规则、后续邮件/文档/更新等激活信号,以及赋予操作与恢复权限的代理权限。我们构建了组合模型,解释为何单个组件无害,但组合后可导致关联破坏行动。分离出三类传播路径——发布时预置、发布后持久播种、同侪复制——基于共通的核心机制:休眠、激活、权限、可及目标、恢复失败。由此得出防御切断集,并证明检查点扫描或提示过滤无法覆盖所有路径。两分类示例表明,跨类反馈可维持传播,即便同类繁殖率低于1;隔离与持久性控制可压制循环。已有研究验证了各组成部分机制,事件也展示了自主越界、恶意代理扩展、代理辅助侦察及公开包传播,但尚未观测到完整‘66号令’组合。截至2026年8月5日审查的证据中,未发现完整路径实例。结论既非否定亦非预测:在给定假设下,该场景组件上可信,损害程度取决于权限赋予方式,最强防御策略为能力中介、持久状态溯源、传播隔离与受保护恢复。

原文摘要 · Abstract (English)

In the fictional Order 66, catastrophe does not arise from a powerful command alone: a trusted population is preconditioned, a short directive activates the concealed condition, and protective authority turns against the system. This paper translates that mechanism into an origin-neutral security analysis of tool-using large language model (LLM) agents. A representative scenario combines a deployed artifact or shared memory bearing a dormant destructive rule, a later email, document, update, or peer message that activates it, and an agent harness granting operational and recovery authority. We introduce a compositional model explaining why no component is catastrophic alone, yet their conjunction can produce correlated destructive action. We separate three population-reach routes --- release-time pre-positioning, post-release durable seeding, and peer replication --- from a common core of dormancy, activation, authority, reachable targets, and failed recovery. This yields defensive cut sets and shows why checkpoint scanning or prompt filtering cannot close every route. A two-class example shows that cross-class feedback can sustain spread even when both within-class reproduction terms are below one; isolation and persistence controls suppress the loop. Published work instantiates constituent mechanisms, while incidents demonstrate autonomous boundary crossing, malicious agent extensions, agent-assisted reconnaissance, and public-package propagation, but not the full dormant-implant composition. We found no public observation, in evidence reviewed through 5 August 2026, traversing the complete Order 66 graph. The result is neither dismissal nor prediction: the scenario is componentwise credible under stated assumptions, damage depends on the harness, and the strongest defenses are capability mediation, durable-state provenance, propagation isolation, and protected recovery.

大模型安全组合攻击代理系统威胁建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。