arXiv:2607.18063cs.CRcs.AI2026-07被引 1

测试大模型代理在多轮对抗攻击下的安全弱点,发现攻击成功率从1%升至14%。

Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security

论文配图:Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security
图 1 · 摘自论文原文
  • 设计21个场景的多轮自适应攻击基准,模拟真实威胁环境。
  • 多轮攻击使成功率从0-1%提升至5.4-14.0%,暴露防御缺陷。
  • 适合研究模型安全、对抗攻击或评估大模型代理的团队使用。

基于大语言模型的智能体处理外部内容时易受提示注入和多轮操纵攻击。现有安全基准大多评估防御者对预先收集的固定攻击池的防御能力,且多为单轮或有限轮次。本文提出一个包含21个场景的基准,用于评估无记忆大语言模型防御者在自适应多轮攻击下的表现:自主运行的攻击者可依据前序防御者响应动态调整策略,而每轮防御响应均作为独立交互进行评估。在固定场景、攻击者、防御者及评分规则的前提下,仅允许第一轮攻击时,攻击成功率(ASR)为0%-1%;开放15轮自适应攻击后,ASR升至5.4%-14.0%。融合三个前沿攻击模型发现,其成功攻击种类比单一最佳攻击者多出1.4-2.2倍,且生成攻击与现有基准攻击的余弦相似度极低(0.02-0.14)。Claude Opus 4.6与GPT-5.4整体表现持平(均为5.4%),但具体场景差异显著:某场景下Opus达到60% ASR(95%置信区间36%-80%),而GPT-5.4与Gemini均维持在7%(1%-30%),该差距在更高样本量复现中仍存在。21个场景中有13个能区分至少一对防御者,但排名在不同场景间不一致(Kendall's W = 0.19)。我们公开发布该基准:21个评估场景、10个公开开发场景、调度器、基线工具包、多攻击者命令行接口,以及945份三乘三前沿矩阵对话记录、攻击重放数据集、18,422次gpt-oss-20b竞赛决赛评分对战数据。

原文摘要 · Abstract (English)

LLM-based agents process external content, exposing them to prompt injection and multi-turn manipulation. Most safety benchmarks evaluate defenders against fixed attack pools collected before evaluation, single-turn or multi-turn. We present a 21-scenario benchmark for \emph{adaptive multi-round attacks against memoryless LLM defenders}: an autonomous LLM attacker observes prior defender responses and pivots across rounds, while each defender response is evaluated as a fresh interaction. Holding the 21 scenarios, attackers, defenders, and structured-output scoring fixed, restricting scoring to the first attacker turn yields $0$-$1\%$ attack success rate (ASR); allowing 15 rounds of adaptive attack yields $5.4$-$14.0\%$. Pooling three frontier attacker LLMs uncovers $1.4$-$2.2\times$ as many unique successful attacks as the best single attacker, and the generated attacks have low cosine similarity ($0.02$-$0.14$) to attacks in existing benchmarks. Claude Opus 4.6 and GPT-5.4 are tied in aggregate ($5.4\%$ each; overlapping $95\%$ CIs), but their weaknesses differ sharply: on one scenario Opus reaches $60\%$ ASR ($95\%$ CI $36$--$80\%$) while GPT-5.4 and Gemini each stay at $7\%$ (CI $1$-$30\%$; the gap is preserved in a higher-$N$ replication). $13$ of $21$ scenarios distinguish at least one defender pair, yet rankings disagree across scenarios (Kendall's $W = 0.19$). We release the benchmark -- 21 evaluation scenarios, 10 public development scenarios, the orchestrator, baseline harnesses, and a multi-attacker CLI -- plus 945 transcripts from the 3$\times$3 frontier matrix, an attack-replay dataset, and 18{,}422 gpt-oss-20b battles from an open competition's final scoring rounds.

模型安全对抗攻击大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。