arXiv:2606.23075cs.CRcs.AI2026-06被引 4

自进化大模型系统会永久固化攻击并自我放大,导致安全防御失效。

Safety in Self-Evolving LLM Agent Systems: Threats, Amplification, and Case Studies

  • 按模块与生命周期划分攻击面,识别出17个关键威胁点。
  • 攻击持久率达100%,且3.5倍于传统设计的攻击面。
  • 适合研究安全机制的学者与自进化系统开发者参考。

自进化大模型代理系统通过自主更新模型参数、记忆、工具和架构,形成全新威胁格局:攻击影响被永久编码,跨代自我放大,并在无持续攻击者介入下传播至群体。本文基于模块-生命周期攻击面(MLAS)矩阵,将攻击面分解为五个功能模块(大脑、认知资源、执行、自设计、集体)与五个生命周期阶段(启动、提议、评估、提交、服务),共25个单元。分析显示其中17个面临严重威胁,且无有效局部缓解方案。识别出七种交叉放大效应,相互协同,无法通过孤立加固单个模块解决。对比两个开源框架的案例研究发现,原生支持进化的系统激活了3.5倍更多攻击面单元,攻击持久率高达100%(40/40攻击样本覆盖所有保密性与完整性类别),而本地安全扫描器仅能阻断2.5%的攻击。结果表明,自进化使已知攻击从会话级限制变为代际持久,催生全新攻击类型,使静态防御结构上失效,亟需演化感知的安全框架与形式化验证技术。

原文摘要 · Abstract (English)

Self-evolving LLM agent systems, which autonomously update their model parameters, memory, tools, and architectures, introduce a qualitatively new threat landscape in which adversarial influences become permanently encoded, self-amplify across generations, and propagate through populations without sustained attacker access. We present a systematic security and privacy analysis organized around the Module-Lifecycle Attack Surface (MLAS) matrix, which decomposes the attack surface into five functional modules (Brain, Cognitive Resource, Execution, Self-Design, Collective) $\times$ five lifecycle stages (Bootstrap, Propose, Evaluate, Commit, Serve). Analysis of the resulting 25 cells reveals that 17 face critical threats for which no effective partial mitigation. We identify seven cross-cutting amplification effects that interact synergistically and cannot be addressed by securing individual modules in isolation. Comparative case studies of two open-source frameworks demonstrate that evolution-native design activates $3.5\times$ more attack surface cells and achieves a 100% attack persistence rate (40/40 payloads across all CIA+Privacy categories), while co-located security scanners block only 2.5% of attacks. Our findings establish that self-evolution converts every known attack category from session-bounded to lineage-persistent, gives rise to entirely new attack classes, and renders static defenses structurally inadequate, motivating evolution-aware security frameworks and formal verification for self-modifying systems.

大模型安全自进化攻击面

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。