让大模型代理自动进化防御能力,应对新型安全威胁。
Beyond Handcrafted Security: Towards Self-Evolving Defense for LLM Agents

- 基于检测机制构建可自主演化的防御框架
- 通过失败日志迭代优化防御策略,提升安全性
- 适合关注自适应安全的大模型应用开发者
大型语言模型(LLM)代理的运行能力持续扩展,带来复杂的安全威胁。运行时防御通过将安全机制嵌入代理执行流程,成为缓解风险的有效手段。然而,现有运行时防御严重依赖人工设计,缺乏系统性构建与维护框架。本文首先提出一种基于检测器(harness)的运行时防御形式化方法,系统刻画检测机制如何支撑防御构建,并统一视角分析现有干预措施。在此基础上,提出HARD(Harness-based Autonomous Runtime Defense Evolution)框架,能够自动识别合适干预策略,并基于观测到的失败轨迹迭代改进防御内容。HARD将防御开发从手工工程转变为自主演化过程,实验表明其在保持正常任务性能的前提下,显著优于传统人工设计的防御方案。研究揭示了自主防御演化作为保障部署中LLM代理安全的新范式,使代理能主动发现弱点并持续强化防护能力。
原文摘要 · Abstract (English)
The expanding operational capabilities of large language model (LLM) agents introduce sophisticated security threats. Runtime defenses have emerged as an effective approach to mitigating these risks by integrating security mechanisms into the agent execution loop. However, existing runtime defenses rely heavily on manually designed interventions and lack a principled framework for their construction and maintenance. In this work, we first develop a harness-level formulation of runtime defense that systematically characterizes how harness mechanisms enable defense construction and provides a unified view of existing runtime defense interventions from a harness perspective. Building on this formulation, we propose HARD (Harness-based Autonomous Runtime Defense Evolution), a self-evolving runtime defense framework that automatically identifies appropriate intervention strategies and iteratively improves defense artifacts based on observed failure traces. HARD transforms runtime defense development from manual engineering into an autonomous evolution process, and extensive experiments demonstrate that it improves security performance over existing handcrafted defenses while preserving benign task utility. Our findings highlight autonomous defense evolution as a promising new paradigm for securing deployed LLM agents, enabling agents to identify defense weaknesses and continuously improve their protection mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。