arXiv:2508.16481cs.LG2025-08被引 12

测试大模型代理系统在恶意攻击下的脆弱性,发现成功率高达90%。

Benchmarking the Robustness of Agentic Systems to Adversarially-Induced Harms

  • 提出新分类体系与基准BAD-ACTS,覆盖多种有害行为。
  • 不同模型受攻击成功率达40%至90%,暴露严重安全隐患。
  • 适合安全研究者、模型开发者关注代理系统的防御设计。

确保代理系统安全使用需全面理解其可能表现出的恶意行为。本文评估基于大语言模型的代理系统在诱发有害行为攻击下的鲁棒性。为此,我们提出一种新的代理系统危害分类法及全新基准BAD-ACTS,用于研究代理系统在多种有害行为下的安全性。BAD-ACTS包含五个不同应用场景的代理系统实现,以及238个高质量有害行为示例和扩展的699个对抗性行为数据集。该基准支持对代理系统在各类有害行为、可用工具及代理间通信结构下的鲁棒性进行全面分析。通过此基准,我们测试了多种攻击者(包括恶意代理和提示注入)引发恶意行为的效果,发现代理系统普遍脆弱,攻击成功率在40%至90%之间。此外,我们提出一种基于零样本消息监控的有效防御方法。我们认为该基准为代理系统安全研究提供了多样化测试环境。代码已公开于https://github.com/JNoether/BAD-ACTS。

原文摘要 · Abstract (English)

Ensuring the safe use of agentic systems requires a thorough understanding of the range of malicious behaviors these systems may exhibit. In this paper, we evaluate the robustness of LLM-based agentic systems against attacks that aim to elicit harmful actions from agents. To this end, we propose a novel taxonomy of harms for agentic systems and a novel benchmark, BAD-ACTS, for studying the security of agentic systems with respect to a wide range of harmful actions. BAD-ACTS consists of five implementations of agentic systems in distinct application environments, as well as a dataset of 238 high-quality examples of harmful actions and an extended dataset containing 699 additional adversarial actions. This enables a comprehensive study of the robustness of agentic systems across a wide range of categories of harmful behaviors, available tools, and inter-agent communication structures. Using this benchmark, we analyze the robustness of agentic systems under an array of attackers attempting to elicit malicious behaviors, including agents acting adversarially and prompt injections. We found that agents are often vulnerable, as indicated by success rates between 40% and 90% depending on the model. We additionally propose an effective defense based on zero-shot message monitoring. We believe that this benchmark provides a diverse testbed for the safety research of agentic systems. Code is available at https://github.com/JNoether/BAD-ACTS.

代理系统安全评测对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。