arXiv:2606.28270cs.AIcs.MA2026-06

为自主智能体设计内生免疫系统,防御运行时攻击。

Agent-Native Immune System: Architecture, Taxonomy, and Engineering

  • 构建六层免疫架构,底层实现物理与逻辑隔离防护。
  • 提出抗病毒与疫苗分类体系,区分表面防御与深层免疫。
  • 支持持续学习的自监控机制,适应新型威胁,适合安全研究者。

从静态聊天机器人向具备持久记忆、工具使用和多智能体协作能力的自主智能体演进,极大地扩展了人工智能威胁面。现有防御机制如边界安全和训练对齐仍处于智能体推理过程之外,难以应对运行时劫持:包括内存污染、工具链操纵和多智能体协议攻击。为此,我们提出首个生物启发的内生防御架构——智能体原生免疫系统(ANIS),嵌入智能体认知循环中。框架包含四项核心贡献:首先,设计六层免疫塔(L0-L5),明确引入非认知性屏障免疫(L1)作为物理与逻辑隔离层;其次,建立智能体病毒与疫苗的统一分类体系,形式化区分非参数化表层防御与参数化稳健疫苗;第三,提出驾驭三元组(元、自、自动)——自监控、元认知自动化核心,驱动持续免疫学习(CIL),使疫苗可动态适应新威胁;最后,确立模型对齐与智能体免疫的理论分界:对齐提供训练阶段的静态价值基础,而ANIS则在运行阶段充当动态“执法”机制。文章结尾提出开放挑战,包括免疫协议标准化、新型评估指标如自免疫率(误报干预率),以及集体智能生态中病原体与疫苗的共演化机制。

原文摘要 · Abstract (English)

The transition from static chat bots to autonomous agents--equipped with persistent memory, tool-use protocols, and multi-agent collaboration--has fundamentally expanded the AI threat landscape. Current defense mechanisms, such as perimeter security and training-time alignment, remain external to the agent's active reasoning loop. Consequently, they fall short: a fully aligned agent remains highly vulnerable to runtime hijacking via memory poisoning, tool-chain manipulation, or multi-agent protocol attacks. To address this critical gap, we introduce the Agent-Native Immune System (ANIS), the first biologically inspired, endogenous defense architecture embedded directly within the agent's cognitive loop. Our framework presents four primary contributions. First, we design a six-layer Immune Tower (L0-L5), distinctly incorporating Barrier Immunity (L1) as a non-cognitive, physical-and-logical isolation layer. Second, we establish a unified taxonomy of Agent Viruses and Agent Vaccines, formalizing the critical distinction between superficial non-parametric defenses and robust parametric vaccines. Third, we conceptualize the Harness Triad--Meta, Self, and Auto--a self-monitoring, meta-cognitive automation backbone that drives Continual Immune Learning (CIL), enabling vaccines to dynamically adapt to novel threats. Finally, we establish a rigorous theoretical demarcation between model alignment and agent immunity: while alignment provides a static "constitutional" value foundation during training, ANIS serves as the dynamic "law enforcement" mechanism during runtime. We conclude by framing open challenges for the field, including immune protocol standardization, novel evaluation metrics such as the Autoimmunity Rate (false-positive intervention rate), and the co-evolutionary dynamics between pathogens and vaccines within collective intelligence ecosystems.

智能体安全免疫系统持续学习防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。