arXiv:2510.13825cs.CRcs.AI2025-10被引 8

为智能体应用打造安全防护层,实现自防御与行为可信

A2AS: Agentic AI Runtime Security and Self-Defense

  • 通过行为证书与认证提示保障运行安全
  • 零延迟、无需重训练,可直接部署于现有系统
  • 适合关注AI安全的开发者与企业用户

A2AS框架作为人工智能代理和大模型应用的安全层,类比于HTTPS对HTTP的保护。该框架通过行为证书强制执行合规行为,激活模型自防御机制,并确保上下文窗口完整性。其定义安全边界,验证提示输入,应用安全规则与定制策略,控制代理行为,构建纵深防御体系。A2AS避免引入延迟开销、外部依赖、架构变更、模型重训及运维复杂性。基础安全模型BASIC构成A2AS核心:(B)行为证书支持行为管控,(A)认证提示保障上下文完整,(S)安全边界实现不可信输入隔离,(I)上下文内防御支持安全推理,(C)编码化策略实现应用定制规则。本文首次提出BASIC模型与A2AS框架,探索其成为行业标准的潜力。

原文摘要 · Abstract (English)

The A2AS framework is introduced as a security layer for AI agents and LLM-powered applications, similar to how HTTPS secures HTTP. A2AS enforces certified behavior, activates model self-defense, and ensures context window integrity. It defines security boundaries, authenticates prompts, applies security rules and custom policies, and controls agentic behavior, enabling a defense-in-depth strategy. The A2AS framework avoids latency overhead, external dependencies, architectural changes, model retraining, and operational complexity. The BASIC security model is introduced as the A2AS foundation: (B) Behavior certificates enable behavior enforcement, (A) Authenticated prompts enable context window integrity, (S) Security boundaries enable untrusted input isolation, (I) In-context defenses enable secure model reasoning, (C) Codified policies enable application-specific rules. This first paper in the series introduces the BASIC security model and the A2AS framework, exploring their potential toward establishing the A2AS industry standard.

AI安全智能体自防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。