arXiv:2603.28166cs.CRcs.AI2026-03中稿 · the FSE 2026 Ideas…被引 5

提出真实工具环境下的代理权限安全评估框架,揭示大模型在权限滥用攻击下的脆弱性。

Evaluating Privilege Usage of Agents with Real-World Tools

  • 构建GrantBox沙箱,自动集成真实工具并模拟真实权限调用
  • 在精心设计攻击下,大模型平均成功率高达84.80%
  • 适用于评估智能体在真实场景中的权限控制能力

赋予大语言模型(LLM)代理真实世界工具可显著提升效率,但同时也将相应权限授予代理与底层模型。不当的权限使用可能导致信息泄露和基础设施破坏。现有安全评测基准多依赖预设工具和受限交互模式,与真实环境差异显著,难以有效评估代理在关键权限控制方面的安全性。为此,本文提出GrantBox——一个用于分析代理权限使用情况的安全评估沙箱。该系统可自动集成真实工具,允许LLM代理调用真实权限,从而在提示注入攻击下评估其权限使用行为。实验结果表明,尽管大模型具备基本安全意识并能阻断部分直接攻击,但仍易受复杂攻击影响,在精心设计场景中平均攻击成功率高达84.80%。

原文摘要 · Abstract (English)

Equipping LLM agents with real-world tools can substantially improve productivity. However, granting agents autonomy over tool use also transfers the associated privileges to both the agent and the underlying LLM. Improper privilege usage may lead to serious consequences, including information leakage and infrastructure damage. While several benchmarks have been built to study agents' security, they often rely on pre-coded tools and restricted interaction patterns. Such crafted environments differ substantially from the real-world, making it hard to assess agents' security capabilities in critical privilege control and usage. Therefore, we propose GrantBox, a security evaluation sandbox for analyzing agent privilege usage. GrantBox automatically integrates real-world tools and allows LLM agents to invoke genuine privileges, enabling the evaluation of privilege usage under prompt injection attacks. Our results indicate that while LLMs exhibit basic security awareness and can block some direct attacks, they remain vulnerable to more sophisticated attacks, resulting in an average attack success rate of 84.80% in carefully crafted scenarios.

大模型安全权限控制智能体评估提示注入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。