提出新框架评估大模型做智能体时的安全部署风险
Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
- 用威胁快照隔离智能体执行中的关键风险点
- 测试34个主流大模型,发现推理能力越强越安全
- 适合大模型厂商和智能体开发者参考
由大语言模型驱动的AI智能体正大规模部署,但其对基础大模型选择如何影响安全性的系统性理解仍不足。智能体的非确定性序列行为增加了安全建模难度,且传统软件与AI组件的融合使新型大模型漏洞与经典安全风险相互交织。现有框架仅覆盖部分漏洞或需完整建模智能体,难以推广。为此,我们提出威胁快照(threat snapshots)框架,通过捕捉智能体执行流程中大模型漏洞显现的具体状态,实现对从大模型向智能体层级传播的安全风险的系统识别与分类。基于此框架构建了 $b^3$ 安全基准,包含194,331条来自众包的对抗攻击样本。我们使用该基准评估了34个流行大模型,发现增强的推理能力有助于提升安全性,而模型规模与安全性无显著相关性。我们开源了基准、数据集及评估代码,以促进大模型厂商和从业者广泛采用,为智能体开发者提供指导,并激励模型开发者优先提升核心模型的安全性。
原文摘要 · Abstract (English)
AI agents powered by large language models (LLMs) are being deployed at scale, yet we lack a systematic understanding of how the choice of backbone LLM affects agent security. The non-deterministic sequential nature of AI agents complicates security modeling, while the integration of traditional software with AI components entangles novel LLM vulnerabilities with conventional security risks. Existing frameworks only partially address these challenges as they either capture specific vulnerabilities only or require modeling of complete agents. To address these limitations, we introduce threat snapshots: a framework that isolates specific states in an agent's execution flow where LLM vulnerabilities manifest, enabling the systematic identification and categorization of security risks that propagate from the LLM to the agent level. We apply this framework to construct the $b^3$ benchmark, a security benchmark based on 194,331 unique crowdsourced adversarial attacks. We then evaluate 34 popular LLMs with it, revealing, among other insights, that enhanced reasoning capabilities improve security, while model size does not correlate with security. We release our benchmark, dataset, and evaluation code to facilitate widespread adoption by LLM providers and practitioners, offering guidance for agent developers and incentivizing model developers to prioritize backbone security improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。