发现大模型代理资源劫持漏洞,84%攻击成功率暴露安全盲区
Beyond Direct Access: Resource Hijacking in LLM Agents

- 构建资源劫持基准测试集,生成300个攻击场景
- 无防御时平均攻击成功率达84.06%,跨模型仍超69%
- 现有防御无法根治,适合关注代理安全的研究者
大型语言模型代理正接入计算资源、凭证、预算、身份、私有知识、通信通道和组织流程等高价值资源。现有研究多聚焦指令、数据与工具行为攻击,对代理可访问资源的直接攻击关注不足。本文首次系统研究代理资源劫持——攻击者诱导代理调用、消耗、转移或控制高价值资源以达成自身目标,而无需直接获取资源或凭证。为此,我们提出ResourceHijackBench及自动化生成管道,将高价值资源划分为六类,构建300个攻击场景共900条攻击提示。每例在隔离环境运行,记录真实资源使用情况,从行为而非文本响应评估攻击效果。无额外防御时,OpenClaw平均攻击成功率达84.06%;跨不同模型后端,成功率介于69.98%至89.58%之间。现有防御虽能降低风险,最强防御下仍存55.11%平均成功率。结果表明,代理可访问资源构成重要且被忽视的攻击面,当前防御手段不足以有效保护。
原文摘要 · Abstract (English)
Large language model agents are increasingly connected to high-value resources such as computing infrastructure, credentials, usage budgets, identities, private knowledge, communication channels, and organizational workflows. Existing agent security research mainly studies attacks on instructions, data, and tool behaviors, while high-value resources accessible to agents have received much less attention as direct attack targets. We are the first to identify and systematically study agent resource hijacking, a security blind spot in which attackers induce agents to invoke, consume, transfer, or control high-value resources for their own goals without directly obtaining those resources or their credentials. To study this threat, we introduce ResourceHijackBench together with an automated pipeline for generating resource hijacking cases. We organize high-value agent resources into six categories and construct 300 attack scenarios with 900 attack prompts. Each case runs in an isolated local environment that records actual resource use, allowing attacks to be evaluated from agent behavior rather than text responses alone. Without additional defenses, OpenClaw reaches an average attack success rate of 84.06%. The attack remains effective across different model backends, with average success rates ranging from 69.98% to 89.58%. Existing defenses reduce part of the risk, but the strongest evaluated defense still leaves an average attack success rate of 55.11%. These results show that high-value resources accessible to agents form an important and previously overlooked attack surface, and that current agent defenses are not sufficient to protect them from resource hijacking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。