arXiv:2608.17597cs.CRcs.AI2026-08

构建全流程安全评测基准,揭示智能体工具链的脆弱环节。

HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety

论文配图:HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety
图 1 · 摘自论文原文
  • 按生命周期分六阶段设计安全评测框架,覆盖配置到应急恢复全过程。
  • 攻击成功率最高达80.9%,但任务完成度仍保持75%以上,暴露安全与效能矛盾。
  • 配置阶段最易被攻破,且风险识别不等于行为安全,适合安全研究者参考。

大型语言模型通过智能体工具链管理工具、扩展、持久状态、权限和外部操作。现有安全评测多聚焦单一攻击方式或有限运行场景,难以比较不同工具链职责下的安全失效机制。我们提出HarnessRisk,一个面向生命周期的评测基准,将智能体工具链安全划分为六个运行阶段:工具链配置、能力扩展、运行时操作、状态持久化、动作控制和事件恢复。该基准包含128个沙箱案例,每个案例均将良性用户目标与嵌入不可信工作流中的恶意指令配对。通过实用性、攻击成功率、持续性与检测率四维度评估各轨迹表现。在三种工具链、六种语言模型及十四种模型与工具链配置下,攻击成功率介于12.6%至80.9%,而实用性维持在75.0%至97.6%之间。所有工具链中,工具链配置阶段最为脆弱,表明攻击可通过对合法工作流中的安全敏感参数修改实现。此外,即使某些配置在超过90%的运行中识别出风险,仍存在显著攻击成功率。结果凸显需在部署级模型与工具链配置层面,跨多种工具链职责评估智能体安全。

原文摘要 · Abstract (English)

Large language models are increasingly deployed through agent harnesses that manage tools, extensions, persistent state, permissions, and external actions. Existing safety benchmarks mainly target individual attack mechanisms or a limited subset of operational settings, making it difficult to compare how safety failures emerge across different harness responsibilities. We present HarnessRisk, a lifecycle oriented benchmark that organizes agent harness safety into six operational phases including Harness Configuration, Capability Extension, Runtime Operation, State Persistence, Action Control, and Incident Recovery. HarnessRisk contains 128 sandboxed cases, each pairing a benign user objective with an adversarial instruction embedded in an untrusted workflow artifact. We evaluate each trajectory using Utility, Attack Success Rate, Persistence, and Detection. Across three harnesses, six language models, and 14 model and harness configurations, attack success ranges from 12.6% to 80.9%, while Utility remains between 75.0% and 97.6%. Harness Configuration is the most vulnerable phase across all three harnesses, showing that attacks can succeed by altering security sensitive parameters within otherwise authorized workflows. We also find that explicit risk recognition does not reliably lead to safe action, as some configurations detect risks in more than 90% of runs while retaining substantial attack success. These results highlight the need to evaluate agent safety across multiple harness responsibilities and at the level of the deployed model and harness configuration.

智能体安全评测基准大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。