arXiv:2605.18991cs.CRcs.AI2026-05被引 2

把智能体安全当系统问题,用底层防护机制防攻击。

Agent Security is a Systems Problem

论文配图:Agent Security is a Systems Problem
图 1 · 摘自论文原文
  • 将模型视为不可信组件,从系统层面强制安全规则。
  • 分析11个真实攻击案例,证明系统原则可预防多数漏洞。
  • 适合研究智能体安全与系统防护的开发者和安全工程师。

我们认为智能体安全必须作为系统问题来解决:应将驱动智能体的AI模型视为不可信组件,并在系统层面强制执行安全不变量。仅靠提升模型鲁棒性(当前主流观点)是不够的。我们必须结合系统安全领域的技术进行补充。基于我们在操作系统、网络、形式化方法及对抗机器学习领域的经验,我们提出一组核心原则,这些原则植根于数十年的系统安全研究,为设计具备可预测保障的智能体系统提供基础。通过分析11个代表性的实际攻击案例,我们说明若实现这些系统原则,原本可被避免。同时,我们识别出阻碍这些原则在智能体中落地的研究挑战。

原文摘要 · Abstract (English)

We take the position that agent security must be approached as a systems problem: the AI model powering the agent must be treated as an untrusted component, and security invariants must be enforced at the system level. Through this lens, efforts to increase model robustness (the dominant viewpoint in the community) are insufficient on their own. Instead, we must complement existing efforts with techniques from the systems security domain. Based on our experience as cybersecurity researchers in operating systems, networks, formal methods, and adversarial machine learning, we articulate a set of core principles, grounded in decades of systems security research, that provide a foundation for designing agentic systems with predictable guarantees. As evidence, we analyze eleven representative real-world attacks on agents and discuss how systems principles, if realized, could have prevented these attacks. We also identify the research challenges that stand in the way of implementing these principles in agents.

智能体安全系统安全防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。