arXiv:2608.06984cs.CRcs.AI2026-08被引 1

评测智能体系统中持久化载体的安全风险传播路径

HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses

论文配图:HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses
图 1 · 摘自论文原文
  • 设计7类持久载体的328个可执行攻击案例,追踪风险传播全生命周期
  • 发现不同载体和模型配置下安全隔离效果差异显著,攻击链进展不一
  • 提出基于执行痕迹的多阶段评估方法,精准定位风险中断点

现代智能体框架通过记忆、技能、工具等持久化载体在任务与会话间保持状态。然而,这种机制带来延迟性安全风险:攻击者注入的内容可跨系统边界留存,并在后续正常请求中触发违规行为。现有基准通常仅覆盖少数载体或框架,且整体攻击成功率无法反映风险传播细节。为此,我们提出HarnessSafe,一个包含328个可执行案例的基准,覆盖七类持久化载体,并在多数主流智能体框架上进行评估。每个案例均以「持久风险生命周期」形式定义,追踪攻击影响从初始注入、跨载体与系统边界的持久存在,到后期良性触发并产生可观测违规的全过程。我们进一步引入基于执行痕迹的多阶段评估方法,通过可观测的执行证据判断每条攻击链的推进程度及终止位置。实验表明,风险隔离效果具有载体特异性,且强烈依赖于框架-模型配置组合。框架与模型后端均显著影响隔离结果,而攻击成功率无法揭示不同生命周期阶段的传播模式差异。

原文摘要 · Abstract (English)

Modern agent harnesses persist state across tasks and sessions through persistent carriers like memory, skills, tools, and shared artifacts. However, this capability creates delayed safety risks: attacker-influenced content can cross system boundaries and later affect the execution of a benign request. Existing benchmarks typically focus on a few carriers or harnesses, while end-to-end attack-success rates reveal little about how risks propagate. To this end, we present HarnessSafe, a benchmark comprising 328 executable cases across seven persistent-carrier families and evaluated on most mainstream agent harnesses. Each case is specified as a Persistent-Risk Lifecycle that traces attacker influence from its initial entry, through persistence across carriers and system boundaries, to a later benign trigger and an observable violation. We further introduce a multi-stage, trace-based evaluation that uses observable execution evidence to determine how far each attack chain progresses and where it is stopped. Experiments show that containment is carrier-specific and strongly depends on the harness-model configuration. Both the harness and model backend substantially shape containment outcomes, while attack success rates cannot reflect distinct lifecycle progression patterns.

智能体安全持久化载体风险传播评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。