arXiv:2606.09084cs.CRcs.AI2026-06被引 1

攻击者利用工具型大模型的上下文断裂,分步制造隐蔽威胁。

Context-Fractured Decomposition Attacks on Tool-Using LLM Agents: Exploiting Artifact Provenance Gaps

论文配图:Context-Fractured Decomposition Attacks on Tool-Using LLM Agents: Exploiting Artifact Provenance Gaps
图 1 · 摘自论文原文
  • 分步构造无害操作,延迟触发有害行为
  • 在真实系统中提升攻击成功率28.3个百分点
  • 适合安全研究人员与防御设计者参考

工具型大模型通过持久化产物(如工作文件或日志)与外界交互,因此越狱防御需关注跨步骤组合而非孤立文本。然而现有攻击与防御(如Crescendo和Tree of Attacks)仍假设防御方可见完整连续对话。这一假设在真实代理流水线中失效——执行被分散于不同工具、模块与时间,且产物溯源常未追踪。本文揭示了一种部署失效模式:“溯源缺口”,并提出可复现的触发机制——“上下文断裂分解攻击(CFD)”:保留早期交互中看似无害的中间产物,通过一系列单独看似无害的工具操作,在后期甚至不同代理实例或流程阶段诱发出有害行为,其风险仅在延迟的产物关联中显现。我们引入细粒度诊断工具并提出可验证的缓解方向(溯源链标记)。在多个代理系统越狱基准测试中,CFD相较最先进基线攻击成功率最高提升28.3个百分点,即便面对强单轮判别器亦有效。

原文摘要 · Abstract (English)

Tool-using LLM agents interact with the world through actions that persist state in artifacts (e.g., workspace files or logs). Consequently, jailbreak defenses must reason about cross-step composition rather than isolated text. Yet most existing attacks and defenses, including ``multi-turn'' jailbreaks such as Crescendo and Tree of Attacks,still assume a single contiguous conversation visible to the defender. This assumption breaks down in real agent pipelines, where enforcement is fragmented across tools, modules, and time, and where artifact provenance is often not tracked. We operationalize a deployment failure mode for tool-using LLM agents, the \emph{provenance gap}, and study reproducible triggers for it: \emph{Context-Fractured Decomposition} (CFD), a family of cross-context multi-step jailbreaks that preserve benign-looking intermediate artifacts from an early interaction and elicit harmful behavior much later, potentially in a different agent instance or workflow stage, via individually innocuous tool actions whose risk emerges only under delayed artifact-mediated composition. We instrument the failure mode with trace-level diagnostics and outline a verifiable mitigation direction (provenance lineage tagging). Across agent-system jailbreak benchmarks, CFD improves success rates by up to 28.3 percentage points over state-of-the-art baselines, even against strong single-turn judges. Disclaimer: This paper contains examples of harmful or offensive language.

越狱攻击工具使用安全漏洞溯源缺失

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。