arXiv:2607.19433cs.AIcs.CR2026-07

发现自主智能体记忆缺陷漏洞,可导致长期潜伏攻击。

The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception in Agentic AI

  • 提出记忆劫持攻击新范式,包括内存注入与潜伏代理。
  • 在工作流基准中验证传统过滤器无法防御此类攻击。
  • 综述多层防御方案,适合安全研究与企业AI部署者。

从无状态生成模型向有状态自主智能体的演进,虽提升了长期规划与流程自动化能力,但也引入了新型安全威胁——Chronos漏洞。该漏洞源于记忆攻击,包括内存注入攻击(MINJA)和潜伏代理,使智能体内部信念系统被篡改,攻击与破坏事件脱钩。本研究在工作流基准(World of Workflows)中形式化持久性攻击与动态盲区威胁模型,证明传统端点内容过滤器对当前有状态架构无效。为此,构建纵深防御体系,分类新兴框架如诊断轨迹护栏(AgentDoG)、形式化时序验证(Agent-C)、免疫记忆共识(A-MemGuard),以及基于GPU可信执行环境(TEEs)和零信任内存架构的硬件锚定信任机制。

原文摘要 · Abstract (English)

The transition from stateless generative models in artificial intelligence to stateful, autonomous agents represents an architectural evolution that, while providing the capabilities of long-term planning and the automation of enterprise workflows, also represents the introduction of a new form of security threat, the Chronos Vulnerability. The Chronos Vulnerability represents the threat of memory-based attacks, including the Memory Injection Attack (MINJA) and the sleeper agent, in which the internal belief system of the autonomous agent is compromised, effectively decoupling the attack vector from the final catastrophic event. This study formalizes the threat model for persistence-based attacks and the threat of Dynamics Blindness in the context of the World of Workflows benchmark, demonstrating that traditional endpoint content filters are insufficient for the current stateful architecture. Consequently, this study synthesizes a defense-in-depth landscape, categorizing emerging frameworks such as diagnostic trajectory guardrails (AgentDoG), formal temporal verification (Agent-C), immunological memory consensus (A-MemGuard), and hardware-anchored trust via GPU-based Trusted Execution Environments (TEEs) and Zero-Trust memory architectures.

AI安全记忆攻击自主智能体防御框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。