arXiv:2506.21967cs.CLcs.LG2025-06被引 2

发现工具集成大模型代理在各环节极易出错,开源模型更脆弱。

More Vulnerable than You Think: On the Stability of Tool-Integrated LLM Agents

  • 测试代理在读文档、选工具、生成参数和处理响应全链条的稳定性
  • 开源模型比闭源模型更易崩溃,模型越大反而越易受骗
  • 适合关注大模型实际应用安全性的研究人员和开发者

当前对工具集成大模型代理的评估多聚焦于端到端工具使用能力,忽视了其稳定性。这限制了其在真实场景的应用,因为内部或外部因素可能导致代理崩溃或异常行为。本研究通过系统实验,考察代理在整个工具调用过程中(包括阅读工具文档、选择工具与生成参数、处理工具响应)是否易受错误影响。结果显示,代理在每个阶段均高度易错;基于开源模型的代理比基于专有模型的更脆弱。此外,增大模型规模并未显著提升工具调用推理能力,反而可能使代理更容易受到伪装成正常指令的攻击。该研究凸显了评估代理稳定性的必要性,为未来大模型开发与评测提供了重要参考。

原文摘要 · Abstract (English)

Current evaluations of tool-integrated LLM agents typically focus on end-to-end tool-usage evaluation while neglecting their stability. This limits their real-world applicability, as various internal or external factors can cause agents to crash or behave abnormally. Our research addresses this by investigating whether agents are vulnerable to errors throughout the entire tool invocation process, including reading tool documentation, selecting tools and generating parameters, and processing the tool's response. Through extensive experiments, we observe that agents are highly susceptible to errors at each stage and agents based on open-source models are more vulnerable than those based on proprietary models. We also find that increasing the model size does not significantly improve tool invocation reasoning and may make agents more vulnerable to attacks resembling normal user instructions. This highlights the importance of evaluating agent stability and offers valuable insights for future LLM development and evaluation.

大模型安全代理系统稳定性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。