arXiv:2607.18826cs.CRcs.AI2026-07中稿 · the Second Worksho…

提出跨代理攻击溯源方法,识别分散在不同会话中的恶意攻击链。

Cross-Agent Campaign Attribution: Linking Asynchronous Attacks Across LLM Agents

论文配图:Cross-Agent Campaign Attribution: Linking Asynchronous Attacks Across LLM Agents
图 1 · 摘自论文原文
  • 设计轻量级指纹向量 $A^2FV$,通过工具使用、时间模式和提示残留判断攻击关联性。
  • 在 SCD-v1 基准上实现 0.82 的成对 AUC,显著优于单会话检测器。
  • 适用于多代理部署场景下的安全评估,尤其适合对抗分布式攻击的防御研究。

LLM 代理防御通常逐会话评估,但实际部署中攻击可能分散在独立代理、团队和运行时之间,导致每个本地防护仅获得零散片段。本文形式化了跨代理异步攻击溯源问题:在无共享运行时状态、测试时攻击标签或攻击者身份信息的前提下,关联来自同一潜在攻击活动的会话。提出轻量级代理端参考协议 Asynchronous Attribution Fingerprint Vectors ($A^2FV$),基于代理可观测的工具使用、时间模式和提示残留,评分成对攻击相似性。构建 SCD-v1 基准,包含良性流量、隔离攻击、多会话攻击链、匹配的非可信逃避与泄漏审计。在 SCD-v1 上,$A^2FV$ 达到 0.82 成对 AUC,而基于分数的单会话检测器与分块大模型裁判仍接近随机水平。结构与风格残留是主要信号源,时间模式则作为诊断通道保留更丰富代理痕迹。交叉风格控制显示信号部分敏感于风格,但不可简化为风格本身。静态与维度感知的非可信压力测试进一步表明,在受控逃避下,成对可区分性依然存在。结果确立了跨代理攻击溯源作为野外部署中保障 LLM 代理安全的独特评估层级。

原文摘要 · Abstract (English)

LLM-agent defenses are typically evaluated one session at a time. In deployment, however, attacks can be distributed across independent agents, teams, and runtimes, leaving each local guardrail with only a sparse fragment. We formalize cross-agent asynchronous campaign attribution: linking sessions from the same latent adversarial campaign without shared runtime state, test-time campaign labels, or attacker identity oracles. We introduce Asynchronous Attribution Fingerprint Vectors ($A^2FV$), a lightweight proxy-side reference protocol for scoring pairwise campaign similarity from proxy-observable tool-use, timing, and prompt residue. We also construct SCD-v1, a controlled persona-matched benchmark with benign traffic, isolated attacks, multi-session campaigns, matched non-oracle evasion, and leakage audits. On SCD-v1, $A^2FV$ achieves 0.82 pairwise AUC for campaign linking, while score-only adaptations of per-session detectors and chunked LLM judges remain near chance under the same task. The strongest fixed signal is carried by structural and stylometric residue, while timing is retained as a diagnostic channel for richer proxy traces. Crossed-style controls show that the signal is partly style-sensitive but not reducible to style alone. Static and dimension-aware non-oracle stress tests further show that pairwise separability persists under controlled evasion. These results establish cross-agent campaign attribution as a distinct evaluation layer for securing LLM agents in the wild.

攻击溯源LLM安全多代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。