为大模型智能体设计分阶段安全防护,提升全流程可靠性。
$S^3$: Improving Agent Safety through Multi-Stage Defense
- 将安全机制抽象为可复用的分阶段技能组件
- 在多阶段风险基准测试中优于现有方法
- 适合构建安全可控的复杂智能体系统
大型语言模型智能体依赖多阶段工作流(如记忆、规划、工具执行)完成复杂任务,但各阶段可能产生风险并跨阶段传播,难以检测与缓解。现有安全方法仅保护孤立阶段,难以集成,缺乏全流程防护。为此,我们提出阶段特定安全技能(Stage-Specific Safety Skills),将异构安全设计统一为具有明确阶段语义的可复用、可组合组件,并建立自动化转换管道将现有安全方案转化为安全技能,构建社区驱动的安全技能库。在此基础上,提出 $S^3$ 多阶段防御框架,由守护代理协调各阶段安全技能,在全流程中实现风险检测与缓解。我们还构建了多阶段风险基准(MSRB)以评估各阶段代表性风险。实验表明,$S^3$ 在安全有效性与任务效用保持方面均持续优于主流基线,验证了阶段特定安全技能作为可扩展、可组合基础,构建鲁棒可信智能体系统的潜力。
原文摘要 · Abstract (English)
Large Language Model (LLM) agents rely on multi-stage agentic workflows, with stages such as memory, planning, and tool execution, to accomplish complex tasks. However, risks may emerge at different stages, propagate across steps, and become difficult to detect and mitigate. Existing safety methods protect only isolated stages and are difficult to integrate, leaving agents without comprehensive protection throughout the workflow. To address these limitations, we introduce Stage-Specific Safety Skills, a unified abstraction that represents heterogeneous safety designs as reusable and composable components with explicit stage semantics. We further develop an automated transformation pipeline that converts existing safety designs into reusable safety skills and establish a community-driven safety skill library. Building on this abstraction, we propose $S^3$, a multi-stage defense framework in which a guard agent orchestrates stage-specific safety skills for risk detection and mitigation throughout the agentic workflow. We also construct the Multi-Stage Risk Benchmark (MSRB) to evaluate representative risks across workflow stages. Experimental results show that $S^3$ consistently outperforms representative state-of-the-art baselines in both safety effectiveness and utility preservation. These results demonstrate the potential of stage-specific safety skills as a scalable and composable foundation for building resilient and trustworthy agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。